pingzh commented on issue #6523: URL: https://github.com/apache/datafusion-comet/issues/6523#issuecomment-5943586619
hi @unikdahal we just had this pr https://github.com/apache/datafusion-comet/pull/6373 Yes, we tested the native writer/reader end-to-end against a real Celeborn cluster, including TPC-H queries with results checked against Spark. This was with our Comet fork and a patched Celeborn 0.6.1 client on Spark 4 / Scala 2.13. The catch is that our fork uses different completion tracking. Those runs don’t validate the stricter checks now in Apache main, and I haven’t verified a client that passes those checks. I’m not aware of a public Celeborn PR providing the completion API we need across retries, cancellation, and transport buffer release. Our fork still relies on private client internals. So for your Spark 3.5 / Celeborn 0.7.0 setup, the fallback is expected. Native shuffle on Apache main still needs compatible completion support. Setting spark.comet.shuffle.mode=native alone won’t enable it. For reference, these are the core settings we use: ``` spark.plugins=org.apache.spark.CometPlugin spark.shuffle.manager=org.apache.spark.sql.comet.execution.shuffle.CometCelebornShuffleManager spark.comet.exec.enabled=true spark.comet.shuffle.enabled=true spark.comet.shuffle.mode=native spark.celeborn.client.spark.stageRerun.enabled=true ``` They need the matching client on both the driver and executors, and won’t address the completion check you’re hitting. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
