https://github.com/aaronj0 created 
https://github.com/llvm/llvm-project/pull/226977

Previously, GeneratePTX() returned an error when PassManager::run() returned 
false. That value reports whether any pass changed the module, not whether 
emission succeeded. IIUC, addPassesToEmitFile() should be the only failure 
point, which is already checked.

The check was harmless until e8b75c172810 ("[NVPTX] Add NewPM boilerplate to 
NVPTXAssignValidGlobalNames"). Before it, that pass returned true 
unconditionally, so run() always reported a change. Now a device module without 
a function definition, produced in the case of host-only inputs, reports no 
change and fails the check, falsely erroring out with `Failed to emit PTX 
code.` Due to this, every host-only input to `clang-repl --cuda` is now 
rejected on main. The host-supports-cuda lit probe consists mostly of such 
inputs, so clang/test/Interpreter/CUDA reports UNSUPPORTED instead of failing.

I've also added a test with host-only inputs so a regression like this could be 
caught.

>From 7f3c292ba208abbc311c1bccfaeba53a9a79e82e Mon Sep 17 00:00:00 2001
From: Aaron Jomy <[email protected]>
Date: Thu, 24 Sep 2026 16:02:54 +0200
Subject: [PATCH] [clang-repl] Fix PTX emission for CUDA inputs with no device
 code

Previously, GeneratePTX() returned an error when PassManager::run()
returned false. That value reports whether any pass changed the module,
not whether emission succeeded. addPassesToEmitFile() should be the only
failure point, which is already checked.

The check was harmless until e8b75c172810 ("[NVPTX] Add NewPM
boilerplate to NVPTXAssignValidGlobalNames"). Before it, that pass
returned true unconditionally, so run() always reported a change. Now a
device module without a function definition, produced in the case of
host-only inputs, reports no change and fails the check, falsely
erroring out with `Failed to emit PTX code.` Due to this, every
host-only input to `clang-repl --cuda` is now rejected on main. The
host-supports-cuda lit probe consists mostly of such inputs, so
clang/test/Interpreter/CUDA reports UNSUPPORTED instead of failing.

I've also added a test with host-only inputs so something like this
could be caught.
---
 clang/lib/Interpreter/DeviceOffload.cpp        |  4 +---
 .../Interpreter/CUDA/empty-device-module.cu    | 18 ++++++++++++++++++
 2 files changed, 19 insertions(+), 3 deletions(-)
 create mode 100644 clang/test/Interpreter/CUDA/empty-device-module.cu

diff --git a/clang/lib/Interpreter/DeviceOffload.cpp 
b/clang/lib/Interpreter/DeviceOffload.cpp
index bf7653c518c30..d876e4cf10c3d 100644
--- a/clang/lib/Interpreter/DeviceOffload.cpp
+++ b/clang/lib/Interpreter/DeviceOffload.cpp
@@ -68,9 +68,7 @@ llvm::Expected<llvm::StringRef> 
IncrementalCUDADeviceParser::GeneratePTX() {
         llvm::inconvertibleErrorCode());
   }
 
-  if (!PM.run(*PTU.TheModule))
-    return llvm::make_error<llvm::StringError>("Failed to emit PTX code.",
-                                               llvm::inconvertibleErrorCode());
+  PM.run(*PTU.TheModule);
 
   PTXCode += '\0';
   while (PTXCode.size() % 8)
diff --git a/clang/test/Interpreter/CUDA/empty-device-module.cu 
b/clang/test/Interpreter/CUDA/empty-device-module.cu
new file mode 100644
index 0000000000000..fd9ba760e5a69
--- /dev/null
+++ b/clang/test/Interpreter/CUDA/empty-device-module.cu
@@ -0,0 +1,18 @@
+// Tests host-only inputs. They produce an empty device module, and emitting
+// PTX for it must not be reported as a failure just because no pass changed
+// the module.
+// RUN: cat %s | clang-repl --cuda | FileCheck %s
+
+extern "C" int printf(const char*, ...);
+
+int host_only = 42;
+printf("host_only: %d\n", host_only);
+// CHECK: host_only: 42
+
+__global__ void kernel() {}
+
+kernel<<<1,1>>>();
+printf("CUDA Error: %d\n", cudaGetLastError());
+// CHECK-NEXT: CUDA Error: 0
+
+%quit

_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits

Reply via email to