https://gcc.gnu.org/bugzilla/show_bug.cgi?id=117439

--- Comment #10 from GCC Commits <cvs-commit at gcc dot gnu.org> ---
The releases/gcc-14 branch has been updated by Jakub Jelinek
<ja...@gcc.gnu.org>:

https://gcc.gnu.org/g:67379c5b6ea4c87d69dea90ede51f33c9f5170c8

commit r14-11135-g67379c5b6ea4c87d69dea90ede51f33c9f5170c8
Author: Jakub Jelinek <ja...@redhat.com>
Date:   Wed Nov 6 10:21:09 2024 +0100

    store-merging: Don't use sub_byte_op_p mode for empty_ctor_p unless
necessary [PR117439]

    encode_tree_to_bitpos uses the more expensive sub_byte_op_p mode in which
    it has to allocate a buffer and do various extra work like shifting the
bits
    etc. if bitlen or bitpos aren't multiples of BITS_PER_UNIT, or if bitlen
    doesn't have corresponding integer mode.
    The last case is explained later in the comments:
      /* The native_encode_expr machinery uses TYPE_MODE to determine how many
         bytes to write.  This means it can write more than
         ROUND_UP (bitlen, BITS_PER_UNIT) / BITS_PER_UNIT bytes (for example
         write 8 bytes for a bitlen of 40).  Skip the bytes that are not within
         bitlen and zero out the bits that are not relevant as well (that may
         contain a sign bit due to sign-extension).  */
    Now, we've later added empty_ctor_p support, either {} CONSTRUCTOR
    or {CLOBBER}, which doesn't use native_encode_expr at all, just memset,
    so that case doesn't need those fancy games unless bitlen or bitpos
    aren't multiples of BITS_PER_UNIT (unlikely, but let's pretend it is
    possible).

    The following patch makes us use the fast path even for empty_ctor_p
    which occupy full bytes, we can just memset that in the provided buffer and
    don't need to XALLOCAVEC another buffer.

    This patch in itself fixes the testcase from the PR (which was about using
    huge XALLLOCAVEC), but I want to do some other changes, to be posted in a
    next patch.

    2024-11-06  Jakub Jelinek  <ja...@redhat.com>

            PR tree-optimization/117439
            * gimple-ssa-store-merging.cc (encode_tree_to_bitpos): For
            empty_ctor_p use !sub_byte_op_p even if bitlen doesn't have an
            integral mode.

    (cherry picked from commit aab572240a0752da74029ed9f8918e0b1628e8ba)
  • [Bug tree-optimization/117439] ... cvs-commit at gcc dot gnu.org via Gcc-bugs

Reply via email to