https://github.com/aaronj0 created
https://github.com/llvm/llvm-project/pull/226977
Previously, GeneratePTX() returned an error when PassManager::run() returned
false. That value reports whether any pass changed the module, not whether
emission succeeded. IIUC, addPassesToEmitFile() should be the only failure
point, which is already checked.
The check was harmless until e8b75c172810 ("[NVPTX] Add NewPM boilerplate to
NVPTXAssignValidGlobalNames"). Before it, that pass returned true
unconditionally, so run() always reported a change. Now a device module without
a function definition, produced in the case of host-only inputs, reports no
change and fails the check, falsely erroring out with `Failed to emit PTX
code.` Due to this, every host-only input to `clang-repl --cuda` is now
rejected on main. The host-supports-cuda lit probe consists mostly of such
inputs, so clang/test/Interpreter/CUDA reports UNSUPPORTED instead of failing.
I've also added a test with host-only inputs so a regression like this could be
caught.
>From 7f3c292ba208abbc311c1bccfaeba53a9a79e82e Mon Sep 17 00:00:00 2001
From: Aaron Jomy <[email protected]>
Date: Thu, 24 Sep 2026 16:02:54 +0200
Subject: [PATCH] [clang-repl] Fix PTX emission for CUDA inputs with no device
code
Previously, GeneratePTX() returned an error when PassManager::run()
returned false. That value reports whether any pass changed the module,
not whether emission succeeded. addPassesToEmitFile() should be the only
failure point, which is already checked.
The check was harmless until e8b75c172810 ("[NVPTX] Add NewPM
boilerplate to NVPTXAssignValidGlobalNames"). Before it, that pass
returned true unconditionally, so run() always reported a change. Now a
device module without a function definition, produced in the case of
host-only inputs, reports no change and fails the check, falsely
erroring out with `Failed to emit PTX code.` Due to this, every
host-only input to `clang-repl --cuda` is now rejected on main. The
host-supports-cuda lit probe consists mostly of such inputs, so
clang/test/Interpreter/CUDA reports UNSUPPORTED instead of failing.
I've also added a test with host-only inputs so something like this
could be caught.
---
clang/lib/Interpreter/DeviceOffload.cpp | 4 +---
.../Interpreter/CUDA/empty-device-module.cu | 18 ++++++++++++++++++
2 files changed, 19 insertions(+), 3 deletions(-)
create mode 100644 clang/test/Interpreter/CUDA/empty-device-module.cu
diff --git a/clang/lib/Interpreter/DeviceOffload.cpp
b/clang/lib/Interpreter/DeviceOffload.cpp
index bf7653c518c30..d876e4cf10c3d 100644
--- a/clang/lib/Interpreter/DeviceOffload.cpp
+++ b/clang/lib/Interpreter/DeviceOffload.cpp
@@ -68,9 +68,7 @@ llvm::Expected<llvm::StringRef>
IncrementalCUDADeviceParser::GeneratePTX() {
llvm::inconvertibleErrorCode());
}
- if (!PM.run(*PTU.TheModule))
- return llvm::make_error<llvm::StringError>("Failed to emit PTX code.",
- llvm::inconvertibleErrorCode());
+ PM.run(*PTU.TheModule);
PTXCode += '\0';
while (PTXCode.size() % 8)
diff --git a/clang/test/Interpreter/CUDA/empty-device-module.cu
b/clang/test/Interpreter/CUDA/empty-device-module.cu
new file mode 100644
index 0000000000000..fd9ba760e5a69
--- /dev/null
+++ b/clang/test/Interpreter/CUDA/empty-device-module.cu
@@ -0,0 +1,18 @@
+// Tests host-only inputs. They produce an empty device module, and emitting
+// PTX for it must not be reported as a failure just because no pass changed
+// the module.
+// RUN: cat %s | clang-repl --cuda | FileCheck %s
+
+extern "C" int printf(const char*, ...);
+
+int host_only = 42;
+printf("host_only: %d\n", host_only);
+// CHECK: host_only: 42
+
+__global__ void kernel() {}
+
+kernel<<<1,1>>>();
+printf("CUDA Error: %d\n", cudaGetLastError());
+// CHECK-NEXT: CUDA Error: 0
+
+%quit
_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits