With Zcmp (cm.push) and -Os, the prologue folds the callee-saved FPRs
(fs0..fs11, i.e. f8/f9/f18..f31) into the stack region that cm.push
reserves using get_multi_push_fpr_mask (multi_push_additional /
UNITS_PER_WORD).  That region is measured in bytes, but a double is
8 bytes (UNITS_PER_FP_REG), so dividing by the word size (4 bytes)
lets the mask fold up to twice as many FPRs as the reserved space can
actually hold.

The surplus FPRs (typically fs6..fs11, i.e. f22..f27) therefore end up
stored at negative offsets of the current SP, with the stack allocation
that should cover them deferred to a trailing "addi sp,sp,-N" that
comes after the stores.  This is a use-the-stack-before-allocating-it
defect: callee-saved FP state lives in the not-yet-allocated region
below SP, which an asynchronous interrupt running on the interrupted
SP (e.g. a nested CLIC/Zcswl handler) can overwrite before the frame is
claimed, silently corrupting f22..f27.

gcc/ChangeLog:

        * config/riscv/riscv.cc (riscv_expand_prologue): Fold into the
        cm.push reserve only as many callee-saved FPRs as it can hold by
        dividing the additional reserve by UNITS_PER_FP_REG.
        (riscv_expand_epilogue): Likewise.

gcc/testsuite/ChangeLog:

        * gcc.target/riscv/zcmp_fpr_save_below_sp.c: New test.  Compile an
        FP-register-heavy function at -Os with Zcmp and check that no
        callee-saved FPR is saved at a negative offset of SP.

Signed-off-by: Ancheng.Qiao <[email protected]>

diff --git a/gcc/config/riscv/riscv.cc b/gcc/config/riscv/riscv.cc
index 9fca2f679a0..b676292ca56 100644
--- a/gcc/config/riscv/riscv.cc
+++ b/gcc/config/riscv/riscv.cc
@@ -10563,7 +10563,7 @@ riscv_expand_prologue (void)
       if (fmask)
        {
          unsigned mask_fprs_push
-           = get_multi_push_fpr_mask (multi_push_additional / UNITS_PER_WORD);
+           = get_multi_push_fpr_mask (multi_push_additional / 
UNITS_PER_FP_REG);
          frame->fmask &= mask_fprs_push;
          riscv_for_each_saved_reg (remaining_size, riscv_save_reg, false,
                                    false, false);
@@ -10969,7 +10969,7 @@ riscv_expand_epilogue (int style)
       if (fmask)
        {
          mask_fprs_push = get_multi_push_fpr_mask (frame->multi_push_adj_addi
-                                                   / UNITS_PER_WORD);
+                                                   / UNITS_PER_FP_REG);
          frame->fmask &= ~mask_fprs_push; /* FPRs not saved by cm.push  */
        }
     }
diff --git a/gcc/testsuite/gcc.target/riscv/zcmp_fpr_save_below_sp.c 
b/gcc/testsuite/gcc.target/riscv/zcmp_fpr_save_below_sp.c
new file mode 100644
index 00000000000..efc629a80d8
--- /dev/null
+++ b/gcc/testsuite/gcc.target/riscv/zcmp_fpr_save_below_sp.c
@@ -0,0 +1,52 @@
+/* { dg-do compile } */
+/* { dg-options "-Os -march=rv32imafd_zca_zcmp -mabi=ilp32d -mcmodel=medlow" } 
*/
+/* { dg-skip-if "" { *-*-* } { "-O0" "-O1" "-O2" "-Og" "-O3" "-Oz" "-flto" } } 
*/
+/* { dg-final { scan-assembler "cm\\.push" } } */
+/* { dg-final { scan-assembler-not "fsd\\s+fs\[0-9\]+,-" } } */
+/*
+   When Zcmp (cm.push) is enabled, GCC may allocate only the GPR save area
+   with cm.push and then save the callee-saved FPRs (fs0..fs11) at *negative*
+   offsets of the current SP, deferring the rest of the frame allocation to a
+   trailing "addi sp,sp,-N".  Those below-SP slots are treated as scratch by
+   the interrupt/preemption machinery (e.g. nested interrupts that run on the
+   interrupted SP), so they can be silently clobbered, corrupting the
+   callee-saved FP registers.  The prologue must grow the frame before
+   emitting any callee-saved FPR store, so no "fsd fsN,-...".  */
+
+double gmem[40];
+
+__attribute__((noinline)) double
+helper (double a, double b)
+{
+  return a * b + gmem[0];
+}
+
+__attribute__((noinline)) double
+sink (double x)
+{
+  gmem[3] = x;
+  return x;
+}
+
+/* Keep many callee-saved FPRs live across calls so that GCC must spill
+   fs2..fs11; this is what triggers the below-SP saves with cm.push.  */
+double
+test_zcmp_fpr_save (double a, double b, double c, double d, double e,
+                   double q, double w, double r, double t, double y,
+                   double u, double i)
+{
+  double x0 = a, x1 = b, x2 = c, x3 = d, x4 = e;
+  double x5 = q, x6 = w, x7 = r, x8 = t, x9 = y;
+
+  double acc = helper (x0, x1) + helper (x2, x3) + helper (x4, x5)
+              + helper (x6, x7) + helper (x8, x9);
+  double acc2 = helper (x0, x2) + helper (x4, x6) + helper (x8, x9)
+               + helper (w, t) + helper (r, y) + helper (u, i);
+  double acc3 = helper (x1, x3) + helper (x5, x7) + helper (x9, y)
+               + helper (q, w) + helper (t, r) + helper (x0, x6);
+
+  gmem[1] = acc;
+  gmem[2] = acc2;
+  gmem[4] = acc3;
+  return sink (acc + acc2 + acc3);
+}
-- 
2.53.0

Reply via email to