andygrove opened a new issue, #6477:
URL: https://github.com/apache/datafusion-comet/issues/6477

   ### Describe the bug
   
   When an array or struct `ORDER BY` key holds a null element or field, an 
aggregate window function over the default `RANGE` frame returns the result for 
the whole partition on that row and on every row after it. DataFusion finds the 
end of a `RANGE` frame by comparing keys with `ScalarValue::partial_cmp`, which 
orders a null element above every other value (Postgres semantics), while the 
sort put that key first, as Spark does. Once the frame search passes that row, 
every later row's frame runs to the end of the partition.
   
   `RANK` and `DENSE_RANK` over the same keys are correct.
   
   ### Steps to reproduce
   
   ```sql
   CREATE TABLE r(id INT, p INT, i INT) USING parquet;
   INSERT INTO r VALUES (1, 1, NULL), (2, 1, 1), (3, 1, 1), (4, 1, 2);
   
   SELECT id, SUM(id) OVER (ORDER BY array(i)) AS running FROM r;
   -- Spark: (1, 1), (2, 6), (3, 6), (4, 10)   Comet: 10 on every row
   
   SELECT id, SUM(id) OVER (PARTITION BY p ORDER BY named_struct('x', i)) AS 
running FROM r;
   -- Spark: (1, 1), (2, 6), (3, 6), (4, 10)   Comet: 10 on every row
   ```
   
   ### Expected behavior
   
   The frame of each row ends at its last peer, as in Spark, so the running 
sums are 1, 6, 6 and 10.
   
   ### Additional context
   
   Found while fixing #5507 in #6475, whose fixture leaves the null row out of 
its running sums to stay clear of this.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to