pingzh opened a new issue, #6555:
URL: https://github.com/apache/datafusion-comet/issues/6555

   ## What is the problem the feature request solves?
   
   Comet does not yet have a dedicated native source for Spark's 
`OneRowRelation`: exactly one row with zero columns. Scalar queries and 
constant branches in larger plans should be able to start native execution 
without depending on Spark to produce that row and convert it to Arrow.
   
   Checked Apache Comet `main` at 
[0fdfaad708d5c82195269134e0621dd959e2c0c9](https://github.com/apache/datafusion-comet/commit/0fdfaad708d5c82195269134e0621dd959e2c0c9):
   
   - The [native operator 
registry](https://github.com/apache/datafusion-comet/blob/0fdfaad708d5c82195269134e0621dd959e2c0c9/spark/src/main/scala/org/apache/comet/rules/CometExecRule.scala#L94-L116)
 and [native plan 
protocol](https://github.com/apache/datafusion-comet/blob/0fdfaad708d5c82195269134e0621dd959e2c0c9/native/proto/src/proto/operator.proto#L49-L80)
 have no dedicated one-row source.
   - The existing [Spark-to-columnar row 
path](https://github.com/apache/datafusion-comet/blob/0fdfaad708d5c82195269134e0621dd959e2c0c9/spark/src/main/scala/org/apache/spark/sql/comet/CometSparkToColumnarExec.scala#L113-L122)
 calls `child.execute()` and converts the resulting Spark rows. It can enable 
native consumers, but Spark still produces the input row.
   
   The main motivation is preserving native execution through larger plans. For 
example, #4949 documents a Spark 4.2 TPC-DS `q77a` plan where unsupported 
one-row branches cause unions and downstream aggregates to fall back.
   
   ## Describe the potential solution
   
   Introduce a native one-row source and connect it to Spark's `OneRowRelation` 
planning paths across supported Spark versions. It should produce one 
zero-column row in one partition, allowing eligible consumers to run in Comet 
with `spark.comet.sparkToColumnar.enabled=false`.
   
   Acceptance criteria:
   
   - Recognize the actual Spark one-row relation; do not reinterpret arbitrary 
`RDDScanExec` inputs as single-row sources.
   - Preserve exactly one row, an empty output schema, and single-partition 
execution. Keep this distinct from an empty relation, which produces zero rows.
   - Assert both native source selection and Spark-compatible results for 
scalar projections, `UNION ALL` with a native input, and aggregate consumers. 
Check row counts and duplicate preservation.
   - Include cases such as `SELECT 42` and `SELECT count(*)`, with the 
Spark-to-columnar bridge disabled. Ensure the tested Spark plan actually 
retains a one-row source.
   - Cover the Spark 4.2 plan shape from #4949, including the intended native 
union/aggregate behavior, and exercise AQE on and off.
   - Preserve expression evaluation semantics and existing fallbacks for 
unsupported consumers.
   
   ## Additional context
   
   - #516 was [closed after scalar queries could use the Spark-to-columnar 
bridge](https://github.com/apache/datafusion-comet/issues/516#issuecomment-4391757030),
 following #2422. This request tracks native row generation itself.
   - #4949 tracks the specific Spark 4.2 union/aggregate regression. A bridge 
fix may address that regression independently; this feature should establish a 
native source that does not require the bridge.
   - This request is based on source inspection and the existing upstream 
reports. The illustrative SQL cases above have not been newly executed as 
reproductions.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to