pingzh opened a new issue, #6555: URL: https://github.com/apache/datafusion-comet/issues/6555
## What is the problem the feature request solves? Comet does not yet have a dedicated native source for Spark's `OneRowRelation`: exactly one row with zero columns. Scalar queries and constant branches in larger plans should be able to start native execution without depending on Spark to produce that row and convert it to Arrow. Checked Apache Comet `main` at [0fdfaad708d5c82195269134e0621dd959e2c0c9](https://github.com/apache/datafusion-comet/commit/0fdfaad708d5c82195269134e0621dd959e2c0c9): - The [native operator registry](https://github.com/apache/datafusion-comet/blob/0fdfaad708d5c82195269134e0621dd959e2c0c9/spark/src/main/scala/org/apache/comet/rules/CometExecRule.scala#L94-L116) and [native plan protocol](https://github.com/apache/datafusion-comet/blob/0fdfaad708d5c82195269134e0621dd959e2c0c9/native/proto/src/proto/operator.proto#L49-L80) have no dedicated one-row source. - The existing [Spark-to-columnar row path](https://github.com/apache/datafusion-comet/blob/0fdfaad708d5c82195269134e0621dd959e2c0c9/spark/src/main/scala/org/apache/spark/sql/comet/CometSparkToColumnarExec.scala#L113-L122) calls `child.execute()` and converts the resulting Spark rows. It can enable native consumers, but Spark still produces the input row. The main motivation is preserving native execution through larger plans. For example, #4949 documents a Spark 4.2 TPC-DS `q77a` plan where unsupported one-row branches cause unions and downstream aggregates to fall back. ## Describe the potential solution Introduce a native one-row source and connect it to Spark's `OneRowRelation` planning paths across supported Spark versions. It should produce one zero-column row in one partition, allowing eligible consumers to run in Comet with `spark.comet.sparkToColumnar.enabled=false`. Acceptance criteria: - Recognize the actual Spark one-row relation; do not reinterpret arbitrary `RDDScanExec` inputs as single-row sources. - Preserve exactly one row, an empty output schema, and single-partition execution. Keep this distinct from an empty relation, which produces zero rows. - Assert both native source selection and Spark-compatible results for scalar projections, `UNION ALL` with a native input, and aggregate consumers. Check row counts and duplicate preservation. - Include cases such as `SELECT 42` and `SELECT count(*)`, with the Spark-to-columnar bridge disabled. Ensure the tested Spark plan actually retains a one-row source. - Cover the Spark 4.2 plan shape from #4949, including the intended native union/aggregate behavior, and exercise AQE on and off. - Preserve expression evaluation semantics and existing fallbacks for unsupported consumers. ## Additional context - #516 was [closed after scalar queries could use the Spark-to-columnar bridge](https://github.com/apache/datafusion-comet/issues/516#issuecomment-4391757030), following #2422. This request tracks native row generation itself. - #4949 tracks the specific Spark 4.2 union/aggregate regression. A bridge fix may address that regression independently; this feature should establish a native source that does not require the bridge. - This request is based on source inspection and the existing upstream reports. The illustrative SQL cases above have not been newly executed as reproductions. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
