Several CI runs failed in the libarena parallel tests with -4 (-EINTR) [1],
which says nothing about what actually went wrong.

Two workers can fail like this:

  worker 1: gives up, e.g. the rendezvous times out, sets test_abort and
            returns its own error (-ETIMEDOUT)
  worker 2: sees test_abort and returns -EINTR

-EINTR only means "someone else already gave up", so it carries no
information. Which of the two gets reported depends on the order
pthread_join() collects them, because

        err = err ?: (long)thread_ret;

keeps the first non-zero value and drops the rest. When the -EINTR worker
comes first, the error describing the actual failure is lost.

Let any other error win over -EINTR, and log each worker that returned an
error.

It is still unclear whether the timeouts come from CI load or from a
problem in the test itself. Report the error accurately first, so the next
failure can be diagnosed.

[1]:
https://github.com/kernel-patches/bpf/actions/runs/29867905253/job/88764463566
https://github.com/kernel-patches/bpf/actions/runs/29878191901/job/88794845824

Signed-off-by: Jiayuan Chen <[email protected]>
---
 .../testing/selftests/bpf/prog_tests/libarena.c | 17 ++++++++++++++++-
 1 file changed, 16 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/bpf/prog_tests/libarena.c 
b/tools/testing/selftests/bpf/prog_tests/libarena.c
index df7e4b8dc394..65743920ff69 100644
--- a/tools/testing/selftests/bpf/prog_tests/libarena.c
+++ b/tools/testing/selftests/bpf/prog_tests/libarena.c
@@ -112,13 +112,28 @@ static int run_libarena_parallel_test_workers(struct 
libarena *skel,
 
 
        for (i = 0; i < nthreads; i++) {
+               int worker_err;
+
                ret = pthread_join(threads[i], &thread_ret);
                if (!ASSERT_OK(ret, "pthread_join")) {
                        err = err ?: ret;
                        continue;
                }
 
-               err = err ?: (long)thread_ret;
+               worker_err = (long)thread_ret;
+               if (!worker_err)
+                       continue;
+
+               /*
+                * A worker that bails out because another one already gave up
+                * reports -EINTR. That is collateral damage, so let any other
+                * error win: it tells us what actually went wrong.
+                */
+               if (!err || err == -EINTR)
+                       err = worker_err;
+
+               fprintf(stdout, "%.*s__%d returned %d\n", (int)prefixlen, name,
+                       i, worker_err);
        }
 
        free(threads);
-- 
2.43.0


Reply via email to