unikdahal opened a new issue, #6523:
URL: https://github.com/apache/datafusion-comet/issues/6523

   ### Context
   
   I'm trying to enable the native Comet -> Celeborn shuffle path from the 1.1 
code.
   
   My setup is:
   
   - Spark 3.5.8
   - Scala 2.12
   - Comet 1.1
   - `CometCelebornShuffleManager`
   - `spark.comet.shuffle.mode=native`
   - Celeborn client: 
`org.apache.celeborn:celeborn-client-spark-3-shaded_2.12:0.7.0`
   
   With this setup, Comet rejects the native shuffle path with:
   
   ```text
   Native Celeborn shuffle requires safely published transport completion 
tracking:
   
org.apache.celeborn.common.network.client.TransportClientFactory.clientBootstraps
   must be a volatile instance field
   ```
   
   From reading the current implementation and the reflection compatibility 
tests, this appears to be intentional.
   
   Released Celeborn 0.6.x / 0.7.x clients are rejected from native push 
admission because the required transport completion fields are not safely 
replaceable.
   
   The current user guide also states that native Celeborn shuffle is 
unavailable with released 0.6.x / 0.7.x clients, while the setup example uses 
the shaded 0.7.0 client for the delegated Spark/Celeborn path.
   
   ### Question
   
   What is the intended way to run and validate the actual native Comet -> 
Celeborn shuffle path today?
   
   Specifically:
   
   1. Was the native Celeborn writer/reader path tested end-to-end against a 
real Celeborn cluster?
   
   2. If yes, which Celeborn client version, branch, or patched build was used?
   
   3. Is there an unreleased or development Celeborn client that satisfies the 
transport completion requirements expected by Comet?
   
   4. Is there ongoing work on the Celeborn side to expose a supported 
completion callback/API so Comet no longer needs to depend on private transport 
internals?
   
   5. If there is currently no compatible Celeborn client, should the native 
Celeborn path be considered infrastructure for a future Celeborn release rather 
than something users can enable end-to-end today?
   
   I'm happy to test this against a real Spark + Celeborn deployment and also 
help with the Celeborn-side work if that is the missing piece.
   
   Thanks!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to