On Tue, 11 Aug 2026, Tamar Christina wrote:

> > -----Original Message-----
> > From: Tamar Christina <[email protected]>
> > Sent: 10 August 2026 15:14
> > To: Richard Biener <[email protected]>; Alfie Richards
> > <[email protected]>
> > Cc: [email protected]; nd <[email protected]>; [email protected]
> > Subject: RE: [patch][vect]: convert do-loop edges properly for
> > LOOP_VINFO_EARLY_BREAKS_VECT_PEELED [PR126301]
> > 
> > > -----Original Message-----
> > > From: Richard Biener <[email protected]>
> > > Sent: 10 August 2026 13:21
> > > To: Alfie Richards <[email protected]>
> > > Cc: Tamar Christina <[email protected]>; [email protected];
> > > nd <[email protected]>; [email protected]
> > > Subject: Re: [patch][vect]: convert do-loop edges properly for
> > > LOOP_VINFO_EARLY_BREAKS_VECT_PEELED [PR126301]
> > >
> > > On Wed, 29 Jul 2026, Alfie Richards wrote:
> > >
> > > > On 29/07/2026 14:46, Tamar Christina wrote:
> > > > > The example
> > > > >
> > > > > char b[100];
> > > > > int c(int a) {
> > > > >    unsigned d = 0;
> > > > >    for (; d < a; ++d)
> > > > >      {
> > > > >        if (b[0] + b[d + 1])
> > > > >          return 0;
> > > > >      }
> > > > >    return 1;
> > > > > }
> > > > >
> > > > > when compiled with a partial masked target, e.g. -march=armv9-a -O3
> > > > > generates
> > > > >
> > > > >          ptrue   p6.b, all
> > > > >          whilelo p15.s, wzr, w3
> > > > >          b       .L5
> > > > > .L11:
> > > > >          ld1b    z31.s, p7/z, [x5, x1]
> > > > >          add     z31.h, z31.h, z30.h
> > > > >          uxth    z31.s, p6/m, z31.s
> > > > >          cmpne   p7.s, p7/z, z31.s, #0
> > > > >          b.any   .L4
> > > > > .L5:
> > > > >          mov     x2, x1
> > > > >          mov     p7.b, p15.b
> > > > >          incw    x1
> > > > >          whilelo p15.s, w1, w3
> > > > >          b.any   .L11
> > > > >          mov     w1, w2
> > > > >          add     x3, x4, :lo12:.LANCHOR0
> > > > >          b       .L8
> > > > >
> > > > > Warped BB rotation due to the incorrect profiles aside (different 
> > > > > problem)
> > > > > this loop will exit one iteration early since the IV check (d < a) is
> > > > > checked
> > > > > as (++d < a).
> > > > >
> > > > > This means that when branched to the scalar code we do 2 vector
> > > iterations
> > > > > in
> > > > > the worst case rather than 1.
> > > > >
> > > > > This is due to the fact that the loop is essentially a do-while loop 
> > > > > coming
> > > > > into
> > > > > the vectorizer.  We detect this loop as a
> > > > > LOOP_VINFO_EARLY_BREAKS_VECT_PEELED
> > > > > because it's the same form as a normal while-do loop but where we
> > picked
> > > a
> > > > > different exit than the loop latch exit.
> > > > >
> > > > > Because of the versioning and other checks in the preheader we know
> > that
> > > if
> > > > > you
> > > > > reach the loop body you'll always do at least 1 iteration of the body 
> > > > > (but
> > > > > possibly never the latch).  i.e. the d != 0 check is always true 
> > > > > should you
> > > > > get
> > > > > to the end of the pre-header.
> > > > >
> > > > > The codegen reflects this but because of the early IV exit check we 
> > > > > never
> > > > > get to
> > > > > the body.
> > > > >
> > > > > Alfie's patch to fix this changed the `d++ < a` check to the correct 
> > > > > (d < a)
> > > > > check, but due to how the control flow becomes we end up doing
> > > > >
> > > > > mask_1 = PHI <mask_0, next_mask>
> > > > > if (mask_1)
> > > > >
> > > > > and causes us to retest the mask, even though the mask generation
> > which
> > > > > happens
> > > > > in the pre-header has guaranteed that for the first iteration it's
> > > > > non-empty.
> > > > >
> > > > > So we lose information because the check in the header is used by both
> > the
> > > > > value
> > > > > from the pre-header and the latch on loop back.
> > > > >
> > > > > Instead of doing that this patch just changed the control flow to 
> > > > > reflect
> > > > > that
> > > > > we can always execute the body at least once.  In effect it moves the 
> > > > > IV
> > > > > latch
> > > > > check to the end.
> > > > >
> > > > > This generates:
> > > > >
> > > > >          ptrue   p6.b, all
> > > > >          whilelo p7.s, wzr, w4
> > > > >          b       .L4
> > > > >          .p2align 2,,3
> > > > > .L11:
> > > > >          whilelo p7.s, w1, w4
> > > > >          b.none  .L15
> > > > > .L4:
> > > > >          mov     x2, x1
> > > > >          incw    x1
> > > > >          ld1b    z31.s, p7/z, [x5, x1]
> > > > >          add     z31.h, z31.h, z30.h
> > > > >          uxth    z31.s, p6/m, z31.s
> > > > >          cmpne   p7.s, p7/z, z31.s, #0
> > > > >          b.none  .L11
> > > > > .L7:
> > > > >          mov     w0, 0
> > > > >          ret
> > > > >
> > > > > and when the profiles are fixed (PR117790) we'd get
> > > > >
> > > > >          ptrue   p6.b, all
> > > > >          whilelo p7.s, wzr, w5
> > > > >          .p2align 5,,15
> > > > > .L4:
> > > > >          mov     x2, x1
> > > > >          incw    x1
> > > > >          ld1b    z31.s, p7/z, [x3, x1]
> > > > >          add     z31.h, z31.h, z30.h
> > > > >          uxth    z31.s, p6/m, z31.s
> > > > >          cmpne   p7.s, p7/z, z31.s, #0
> > > > >          b.any   .L7
> > > > >          whilelo p7.s, w1, w5
> > > > >          b.any   .L4
> > > > >
> > > > > Bootstrapped Regtested on aarch64-none-linux-gnu,
> > > > > arm-none-linux-gnueabihf, x86_64-pc-linux-gnu
> > > > > -m32, -m64 and no issues.
> > > > >
> > > > > Any comments?
> > > >
> > > > FWIW LGTM, fixes the issues I was having and markedly improved code gen
> > > for
> > > > the cases I was looking at.
> > > >
> > > > There are some minor extra changes needed for FFR but I will submit 
> > > > those
> > > with
> > > > FFR series and related changes once I finish squashing bugs.
> > >
> > > Sorry for chiming in late - so how can this be valid without actually
> > > peeling up to the IV exit, that is, actually performing loop-ch with
> > > the "correct" exit/latch in mind?  And why would that be not
> > > valid for vect_partial_vectors_avx512?  IMO the terminology in the
> > > added comments is confusing.  We have
> > > LOOP_VINFO_EARLY_BREAKS_VECT_PEELED
> > > also for two early exits and the order
> > >
> > >   if (early)
> > >
> > >   if (IV-exit)
> > >
> > >   if (early)
> > >
> > > no?
> > 
> > Yes, but LOOP_VINFO_EARLY_BREAKS_VECT_PEELED always replay
> > the loop from the start of the current vector iteration, and it won't
> > have any side effects.
> > 
> > So when you exit you will always restart the scalar loop.
> > 
> > When this isn't the case you need more updates, or peeling up to the IV
> > exit as you mentioned.
> > 
> > The other exception is UNCOUNTED loops, but uncounted loops today
> > can't use masking because masking atm needs a counted exit.
> > 
> > > And why would that be not valid for vect_partial_vectors_avx512?
> > 
> > I looked at the codegen for both AVX and RISCV for this before, in both
> > cases they seem to be able to generate good code already since the
> > loop control code is a bit different.  As I tried to explain the majority
> > of the issues here is because using loop masks as control mechanisms
> > are expensive.
> > 
> > If concerned I'd suggest just tightening and perhaps renaming
> > vect_use_loop_latch_condition_p to check for this specific condition.
> > 
> > But that means this loop https://godbolt.org/z/4z6hK7szY will stay
> > suboptimal and throwing away expensive mask comparisons (in L14
> > the result of the mask generation is not used at all this iteration for the
> > block.)
> > 
> > Since I don't think the peeling of a partial loop is going to be easy.
> 
> Ok looks like as long as the region is identified, 
> gimple_duplicate_seme_region
> makes it rather simple... at least at first glance..
> 
> I can do this if you prefer.

Well, let's see.  If we can elide all _PEELED variants we could ensure
this in the pre-vectorizer loop header copying pass by rotating the
loop?  Just copying parts does not perform the rotation.

But I think I am not fully getting the problem you are solving :/

Richard.

> 
> Tamar
> > 
> > Thanks,
> > Tamar
> > 
> > >
> > > Richard.
> > >
> > > > >
> > > > > Thanks,
> > > > > Tamar
> > > > >
> > > > > gcc/ChangeLog:
> > > > >
> > > > >  * tree-vect-loop-manip.cc (vect_use_loop_latch_condition_p): New.
> > > > >  (vect_set_loop_condition_partial_vectors, vect_set_loop_condition):
> > Use
> > > > >  it to rewrite latch and condition.
> > > > >  (vect_set_loop_controls_directly): Adjust IV for do-loop conversion.
> > > > >
> > > > > gcc/testsuite/ChangeLog:
> > > > >
> > > > >  * gcc.target/aarch64/sve/peeled.c: New test.
> > > > >  * gcc.target/aarch64/sve/peeled1.c: New test.
> > > > >  * gcc.target/aarch64/sve/peeled1_run.c: New test.
> > > > >  * gcc.target/aarch64/sve/peeled2.c: New test.
> > > > >  * gcc.target/aarch64/sve/peeled2_run.c: New test.
> > > > >  * gcc.target/aarch64/sve/peeled_run.c: New test.
> > > > >
> > > > > ---
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..f40ffc8f0bcde40aeb75
> > > 4887ed6d5047e9806bc7
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > > @@ -0,0 +1,20 @@
> > > > > +/* { dg-do compile } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +char b[100];
> > > > > +
> > > > > +int __attribute__ ((noipa))
> > > > > +c (int a)
> > > > > +{
> > > > > +  unsigned d = 0;
> > > > > +  for (; d < a; ++d)
> > > > > +    {
> > > > > +      if (b[0] + b[d + 1])
> > > > > +     return 0;
> > > > > +    }
> > > > > +  return 1;
> > > > > +}
> > > > > +
> > > > > +/* { dg-final { scan-assembler-times {\twhilelo\t} 2 } } */
> > > > > +/* { dg-final { scan-assembler-times {\tptest\t} 0 } } */
> > > > > +/* { dg-final { scan-assembler {\tld1b\tz[0-9]+\.s, p[0-9]+/z,} } } 
> > > > > */
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..92e5bca44c48cefa20e
> > > 829510a8914e4852145ed
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > > @@ -0,0 +1,24 @@
> > > > > +/* { dg-do compile } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +char b[100];
> > > > > +char e[100];
> > > > > +
> > > > > +int __attribute__ ((noipa))
> > > > > +c (int a)
> > > > > +{
> > > > > +  unsigned d = 0;
> > > > > +  for (; d < a; ++d)
> > > > > +    {
> > > > > +      if (b[0] + b[d + 1])
> > > > > +     return 0;
> > > > > +
> > > > > +      if (e[0] + e[d + 1])
> > > > > +     return 0;
> > > > > +    }
> > > > > +  return 1;
> > > > > +}
> > > > > +
> > > > > +/* { dg-final { scan-assembler-times {\twhilelo\t} 2 } } */
> > > > > +/* { dg-final { scan-assembler-times {\tptest\t} 2 } } */
> > > > > +/* { dg-final { scan-assembler-times {\tld1b\tz[0-9]+\.h, 
> > > > > p[0-9]+/z,} 2 } }
> > > > > */
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled1_run.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled1_run.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..8ec95a187ab0bce43d
> > > 7bcf9e8d83ea8ddae7ccf5
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled1_run.c
> > > > > @@ -0,0 +1,34 @@
> > > > > +/* { dg-do run { target aarch64_sve_hw } } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +#include "peeled1.c"
> > > > > +
> > > > > +static void
> > > > > +clear_arrays (void)
> > > > > +{
> > > > > +  for (int i = 0; i < 100; ++i)
> > > > > +    {
> > > > > +      b[i] = 0;
> > > > > +      e[i] = 0;
> > > > > +    }
> > > > > +}
> > > > > +
> > > > > +int
> > > > > +main (void)
> > > > > +{
> > > > > +  clear_arrays ();
> > > > > +  if (c (99) != 1)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  clear_arrays ();
> > > > > +  b[37] = 1;
> > > > > +  if (c (99) != 0)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  clear_arrays ();
> > > > > +  e[45] = 1;
> > > > > +  if (c (99) != 0)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  return 0;
> > > > > +}
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled2.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled2.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..c29dab4b5536d318b4
> > > 5a0b70c407ef4d9fa32ed5
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled2.c
> > > > > @@ -0,0 +1,22 @@
> > > > > +/* { dg-do compile } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +char b[100];
> > > > > +
> > > > > +int __attribute__ ((noipa))
> > > > > +c (int a)
> > > > > +{
> > > > > +  unsigned d = 0;
> > > > > +  do
> > > > > +    {
> > > > > +      if (b[0] + b[d + 1])
> > > > > +     return 0;
> > > > > +      d++;
> > > > > +    }
> > > > > +  while (__builtin_expect (d < a, 1));
> > > > > +  return 1;
> > > > > +}
> > > > > +
> > > > > +/* { dg-final { scan-assembler-times {\twhilelo\t} 2 } } */
> > > > > +/* { dg-final { scan-assembler-times {\tptest\t} 0 } } */
> > > > > +/* { dg-final { scan-assembler {\tld1b\tz[0-9]+\.s, p[0-9]+/z,} } } 
> > > > > */
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled2_run.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled2_run.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..2075b0803eb0cab439
> > > 32bc9e60f80f0822ffa83b
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled2_run.c
> > > > > @@ -0,0 +1,26 @@
> > > > > +/* { dg-do run { target aarch64_sve_hw } } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +#include "peeled2.c"
> > > > > +
> > > > > +static void
> > > > > +clear_b (void)
> > > > > +{
> > > > > +  for (int i = 0; i < 100; ++i)
> > > > > +    b[i] = 0;
> > > > > +}
> > > > > +
> > > > > +int
> > > > > +main (void)
> > > > > +{
> > > > > +  clear_b ();
> > > > > +  if (c (99) != 1)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  clear_b ();
> > > > > +  b[37] = 1;
> > > > > +  if (c (99) != 0)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  return 0;
> > > > > +}
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled3.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled3.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..8d82828b770d16868c
> > > 99930c115571c9128c3fa5
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled3.c
> > > > > @@ -0,0 +1,21 @@
> > > > > +/* { dg-do compile } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable -fdump-tree-vect-details" } */
> > > > > +
> > > > > +int __attribute__ ((noipa))
> > > > > +c (int *restrict x, int *restrict y, int n)
> > > > > +{
> > > > > +  unsigned d = 5;
> > > > > +  for (; __builtin_expect (d < n, 1); ++d)
> > > > > +    {
> > > > > +      if (x[d] != y[d])
> > > > > +     return 0;
> > > > > +    }
> > > > > +  return 1;
> > > > > +}
> > > > > +
> > > > > +/* { dg-final { scan-tree-dump "Both peeling and versioning will be
> > > > > applied" "vect" } } */
> > > > > +/* { dg-final { scan-tree-dump "misalignment for fully-masked loop"
> > > "vect"
> > > > > } } */
> > > > > +/* { dg-final { scan-assembler {\tsub\tw[0-9]+, w[0-9]+, #6} } } */
> > > > > +/* { dg-final { scan-assembler-times {\twhilelo\t} 3 } } */
> > > > > +/* { dg-final { scan-assembler-times {\tptest\t} 0 } } */
> > > > > +/* { dg-final { scan-assembler-times {\tld1w\t} 2 } } */
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled3_run.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled3_run.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..83015b3bfe1861e951
> > > 71f59f4a8ddd679b9583d2
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled3_run.c
> > > > > @@ -0,0 +1,35 @@
> > > > > +/* { dg-do run { target aarch64_sve_hw } } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +#include "peeled3.c"
> > > > > +
> > > > > +#define N 128
> > > > > +
> > > > > +int a[N] __attribute__ ((aligned (64)));
> > > > > +int b[N] __attribute__ ((aligned (64)));
> > > > > +
> > > > > +static void
> > > > > +clear_arrays (void)
> > > > > +{
> > > > > +  for (int i = 0; i < N; ++i)
> > > > > +    {
> > > > > +      a[i] = 0;
> > > > > +      b[i] = 0;
> > > > > +    }
> > > > > +}
> > > > > +
> > > > > +int
> > > > > +main (void)
> > > > > +{
> > > > > +  clear_arrays ();
> > > > > +  b[16] = 1;
> > > > > +  if (c (a, b, 16) != 1)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  clear_arrays ();
> > > > > +  b[15] = 1;
> > > > > +  if (c (a, b, 16) != 0)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  return 0;
> > > > > +}
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled_run.c
> > > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled_run.c
> > > > > new file mode 100644
> > > > > index
> > > > >
> > >
> > 0000000000000000000000000000000000000000..7e289f3c0c29fda37f3
> > > 138f3655e79389d76a2d0
> > > > > --- /dev/null
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled_run.c
> > > > > @@ -0,0 +1,26 @@
> > > > > +/* { dg-do run { target aarch64_sve_hw } } */
> > > > > +/* { dg-options "-O3 -mautovec-preference=sve-only
> > > > > -msve-vector-bits=scalable" } */
> > > > > +
> > > > > +#include "peeled.c"
> > > > > +
> > > > > +static void
> > > > > +clear_b (void)
> > > > > +{
> > > > > +  for (int i = 0; i < 100; ++i)
> > > > > +    b[i] = 0;
> > > > > +}
> > > > > +
> > > > > +int
> > > > > +main (void)
> > > > > +{
> > > > > +  clear_b ();
> > > > > +  if (c (99) != 1)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  clear_b ();
> > > > > +  b[37] = 1;
> > > > > +  if (c (99) != 0)
> > > > > +    __builtin_abort ();
> > > > > +
> > > > > +  return 0;
> > > > > +}
> > > > > diff --git a/gcc/tree-vect-loop-manip.cc b/gcc/tree-vect-loop-manip.cc
> > > > > index
> > > > >
> > >
> > ec9dec1b7135916ea69fd9eabfdfe1645d6ea11b..96c10340efd94e7db503af
> > > 208ca21cf3fd807dcd
> > > > > 100644
> > > > > --- a/gcc/tree-vect-loop-manip.cc
> > > > > +++ b/gcc/tree-vect-loop-manip.cc
> > > > > @@ -464,6 +464,19 @@ vect_iv_increment_position (edge loop_exit,
> > > > > @@ gimple_stmt_iterator *bsi,
> > > > >     *insert_after = false;
> > > > >   }
> > > > >
> > > > > +/* If this is a loop where the latch condition should be rewritten to
> > > > > reflect
> > > > > +   a control flow change from a while-do to a do-while loop.  */
> > > > > +
> > > > > +static bool
> > > > > +vect_use_loop_latch_condition_p (loop_vec_info loop_vinfo)
> > > > > +{
> > > > > +  return (loop_vinfo
> > > > > +       && LOOP_VINFO_EARLY_BREAKS_VECT_PEELED (loop_vinfo)
> > > > > +       && LOOP_VINFO_USING_PARTIAL_VECTORS_P (loop_vinfo)
> > > > > +       && (LOOP_VINFO_PARTIAL_VECTORS_STYLE (loop_vinfo)
> > > > > +           != vect_partial_vectors_avx512));
> > > > > +}
> > > > > +
> > > > >   /* Helper for vect_set_loop_condition_partial_vectors.  Generate
> > > > >   definitions
> > > > >      for all the rgroup controls in RGC and return a control that is 
> > > > > nonzero
> > > > >      when the loop needs to iterate.  Add any new preheader 
> > > > > statements to
> > > > > @@ -744,6 +757,13 @@ vect_set_loop_controls_directly (class loop
> > > *loop,
> > > > > @@ loop_vec_info loop_vinfo,
> > > > >                                 bias_tree);
> > > > >    }
> > > > >   +      /* A do-while loop always executes the body once, as such 
> > > > > the limit
> > > > > +      the end counter should be lowered by 1 iteration.  */
> > > > > +      if (vect_use_loop_latch_condition_p (loop_vinfo))
> > > > > +     this_test_limit = gimple_build (preheader_seq, MINUS_EXPR,
> > > > > +                                     compare_type, this_test_limit,
> > > > > +                                     build_one_cst (compare_type));
> > > > > +
> > > > >           /* Create the initial control.  First include all items that
> > > > >     are within the loop limit.  */
> > > > >         tree init_ctrl = NULL_TREE;
> > > > > @@ -968,7 +988,62 @@ vect_set_loop_condition_partial_vectors (class
> > > loop
> > > > > @@ *loop, edge exit_edge,
> > > > >          cond_stmt
> > > > >    = gimple_build_cond (code, test_ctrl, zero_ctrl, NULL_TREE,
> > NULL_TREE);
> > > > >       }
> > > > > -  gsi_insert_before (&loop_cond_gsi, cond_stmt, GSI_SAME_STMT);
> > > > > +  edge latch_exit_edge = NULL;
> > > > > +  /* Convert the loop into a do-while form similar to what ch_vect 
> > > > > would
> > > > > have
> > > > > +     done.  We know that after the checks and peeling that we have at
> > least
> > > > > one
> > > > > +     iteration to perform of the loop because the loop is PEELED.  A 
> > > > > PEELED
> > > > > loop
> > > > > +     has the increment exit before the early ones, i.e. it's a 
> > > > > do-while
> > > > > loop but
> > > > > +     if we materialize the IV edge in that place we are essentially
> > > > > checking one
> > > > > +     iteration ahead so we exit early.  Instead when using masks and 
> > > > > the
> > > > > loop
> > > > > +     is PEELED we remove the existing loop latch and make it a fall 
> > > > > through
> > > > > +     edge and place the latch back to the end of the loop.  So 
> > > > > effectively
> > > > > +     transform:
> > > > > +
> > > > > +     header
> > > > > +       |
> > > > > +     latch
> > > > > +       |
> > > > > +     body
> > > > > +       |
> > > > > +     branch to header
> > > > > +
> > > > > +     into
> > > > > +
> > > > > +     header
> > > > > +       |
> > > > > +     body
> > > > > +       |
> > > > > +     newlatch
> > > > > +       |
> > > > > +     branch to header
> > > > > +
> > > > > +     because the conditions in the pre-header makes it safe to do so 
> > > > > for
> > > > > some
> > > > > +     cases.  */
> > > > > +  if (vect_use_loop_latch_condition_p (loop_vinfo))
> > > > > +    {
> > > > > +      basic_block latch = loop->latch;
> > > > > +      edge latch_e = single_succ_edge (latch);
> > > > > +      int exit_flags = exit_edge->flags & (EDGE_TRUE_VALUE |
> > > > > EDGE_FALSE_VALUE);
> > > > > +
> > > > > +      latch_e->flags &= ~(EDGE_FALLTHRU | EDGE_TRUE_VALUE |
> > > > > EDGE_FALSE_VALUE);
> > > > > +      latch_e->flags |= (EDGE_TRUE_VALUE | EDGE_FALSE_VALUE) ^
> > > exit_flags;
> > > > > +      latch_exit_edge = make_edge (latch, exit_edge->dest, 
> > > > > exit_flags);
> > > > > +      latch_exit_edge->probability = exit_edge->probability;
> > > > > +      latch_exit_edge->count () = exit_edge->count ();
> > > > > +      copy_phi_arg_into_existing_phi (exit_edge, latch_exit_edge);
> > > > > +      gimple_stmt_iterator latch_gsi = gsi_last_bb (latch);
> > > > > +      gsi_insert_after (&latch_gsi, cond_stmt, GSI_NEW_STMT);
> > > > > +      LOOP_VINFO_MAIN_EXIT (loop_vinfo) = latch_exit_edge;
> > > > > +
> > > > > +      gcond *old_cond = as_a <gcond *> (gsi_stmt (loop_cond_gsi));
> > > > > +      if (exit_edge->flags & EDGE_TRUE_VALUE)
> > > > > +     gimple_cond_make_false (old_cond);
> > > > > +      else
> > > > > +     gimple_cond_make_true (old_cond);
> > > > > +      update_stmt (old_cond);
> > > > > +    }
> > > > > +  else
> > > > > +    gsi_insert_before (&loop_cond_gsi, cond_stmt, GSI_SAME_STMT);
> > > > >
> > > > >     /* The loop iterates (NITERS - 1) / VF + 1 times.
> > > > >        Subtract one from this to get the latch count.  */
> > > > > @@ -993,7 +1068,7 @@ vect_set_loop_condition_partial_vectors (class
> > > loop
> > > > > @@ *loop, edge exit_edge,
> > > > >    }
> > > > >           else
> > > > >       assign = gimple_build_assign (final_iv, orig_niters);
> > > > > -      gsi_insert_on_edge_immediate (exit_edge, assign);
> > > > > +      gsi_insert_on_edge_immediate (LOOP_VINFO_MAIN_EXIT
> > > (loop_vinfo),
> > > > > assign);
> > > > >       }
> > > > >
> > > > >     return cond_stmt;
> > > > > @@ -1470,11 +1545,14 @@ vect_set_loop_condition (class loop *loop,
> > > edge
> > > > > @@ loop_e, loop_vec_info loop_vinfo
> > > > >
> > > > >     /* Remove old loop exit test.  */
> > > > >     stmt_vec_info orig_cond_info;
> > > > > -  if (loop_vinfo
> > > > > -      && (orig_cond_info = loop_vinfo->lookup_stmt (orig_cond)))
> > > > > -    loop_vinfo->remove_stmt (orig_cond_info);
> > > > > -  else
> > > > > -    gsi_remove (&loop_cond_gsi, true);
> > > > > +  if (!vect_use_loop_latch_condition_p (loop_vinfo))
> > > > > +    {
> > > > > +      if (loop_vinfo
> > > > > +       && (orig_cond_info = loop_vinfo->lookup_stmt (orig_cond)))
> > > > > +     loop_vinfo->remove_stmt (orig_cond_info);
> > > > > +      else
> > > > > +     gsi_remove (&loop_cond_gsi, true);
> > > > > +    }
> > > > >
> > > > >     if (dump_enabled_p ())
> > > > >       dump_printf_loc (MSG_NOTE, vect_location, "New loop exit
> > condition:
> > > > >       %G",
> > > > >
> > > > >
> > > >
> > > >
> > >
> > > --
> > > Richard Biener <[email protected]>
> > > SUSE Software Solutions Germany GmbH,
> > > Frankenstrasse 146, 90461 Nuernberg, Germany;
> > > GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG
> > > Nuernberg)
> 

-- 
Richard Biener <[email protected]>
SUSE Software Solutions Germany GmbH,
Frankenstrasse 146, 90461 Nuernberg, Germany;
GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG Nuernberg)

Reply via email to