[ 
https://issues.apache.org/jira/browse/FLINK-40919?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Yaoxuan Wu updated FLINK-40919:
-------------------------------
    Description: 
{{AVG}} on an INT column is computed without overflow on its own. When the 
planner rewrites it into {{{}SUM / COUNT{}}}, the sum keeps type INT and wraps 
around, so the same AVG over the same rows returns a wrong value.

Trigger 1: SUM and COUNT of the same column in the SELECT list:

 
{code:java}
SELECT AVG(x) FROM (VALUES (2147483647), (2147483647), (2147483647)) AS v(x);
-- 2147483647 (correct)

SELECT AVG(x), SUM(x), COUNT(x) FROM (VALUES (2147483647), (2147483647), 
(2147483647)) AS v(x);
-- AVG = 715827881 (wrong) {code}
 

 

Trigger 2: AVG over a join, where AVG is split into SUM/COUNT and pushed below 
the join:

 
{code:java}
from pyflink.table import EnvironmentSettings, TableEnvironment, DataTypest_env 
= TableEnvironment.create(EnvironmentSettings.in_batch_mode())

a = t_env.from_elements([(2147483647,), (2147483647,), (1,)], 
DataTypes.ROW([DataTypes.FIELD('x', DataTypes.INT())]))
b = t_env.from_elements([(1,), (2,)], DataTypes.ROW([DataTypes.FIELD('y', 
DataTypes.INT())]))
t_env.create_temporary_view('a', a)
t_env.create_temporary_view('b', b)

# Expected: 1431655765 = (2147483647 + 2147483647 + 1) / 3, each row repeated 
twice by the join
t_env.execute_sql("SELECT AVG(x) FROM a").print()                    # 
1431655765 (correct)
t_env.execute_sql("SELECT AVG(x) FROM a, b").print()                 # 0        
  (wrong)
t_env.execute_sql("SELECT AVG(CAST(x AS BIGINT)) FROM a, b").print() # 
1431655765 (correct)
print(t_env.explain_sql("SELECT AVG(x) FROM a, b"))
{code}
 

  was:
{{AVG(x)}} on an INT column is computed without overflow on its own. When the 
planner rewrites it into {{{}SUM(x) / COUNT(x){}}}, the sum keeps type INT and 
wraps around, so the same AVG over the same rows returns a wrong value.

Trigger 1: SUM and COUNT of the same column in the SELECT list:

 
{code:java}
SELECT AVG(x) FROM (VALUES (2147483647), (2147483647), (2147483647)) AS v(x);
-- 2147483647 (correct)

SELECT AVG(x), SUM(x), COUNT(x) FROM (VALUES (2147483647), (2147483647), 
(2147483647)) AS v(x);
-- AVG = 715827881 (wrong) {code}
 

 

Trigger 2: AVG over a join, where AVG is split into SUM/COUNT and pushed below 
the join:

 
{code:java}
from pyflink.table import EnvironmentSettings, TableEnvironment, DataTypest_env 
= TableEnvironment.create(EnvironmentSettings.in_batch_mode())

a = t_env.from_elements([(2147483647,), (2147483647,), (1,)], 
DataTypes.ROW([DataTypes.FIELD('x', DataTypes.INT())]))
b = t_env.from_elements([(1,), (2,)], DataTypes.ROW([DataTypes.FIELD('y', 
DataTypes.INT())]))
t_env.create_temporary_view('a', a)
t_env.create_temporary_view('b', b)

# Expected: 1431655765 = (2147483647 + 2147483647 + 1) / 3, each row repeated 
twice by the join
t_env.execute_sql("SELECT AVG(x) FROM a").print()                    # 
1431655765 (correct)
t_env.execute_sql("SELECT AVG(x) FROM a, b").print()                 # 0        
  (wrong)
t_env.execute_sql("SELECT AVG(CAST(x AS BIGINT)) FROM a, b").print() # 
1431655765 (correct)
print(t_env.explain_sql("SELECT AVG(x) FROM a, b"))
{code}
 


> AVG on INT returns wrong result when the planner rewrites it into SUM/COUNT 
> (INT sum overflows)
> -----------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40919
>                 URL: https://issues.apache.org/jira/browse/FLINK-40919
>             Project: Flink
>          Issue Type: Bug
>          Components: Table SQL / Planner
>    Affects Versions: 2.3.0
>            Reporter: Yaoxuan Wu
>            Priority: Major
>
> {{AVG}} on an INT column is computed without overflow on its own. When the 
> planner rewrites it into {{{}SUM / COUNT{}}}, the sum keeps type INT and 
> wraps around, so the same AVG over the same rows returns a wrong value.
> Trigger 1: SUM and COUNT of the same column in the SELECT list:
>  
> {code:java}
> SELECT AVG(x) FROM (VALUES (2147483647), (2147483647), (2147483647)) AS v(x);
> -- 2147483647 (correct)
> SELECT AVG(x), SUM(x), COUNT(x) FROM (VALUES (2147483647), (2147483647), 
> (2147483647)) AS v(x);
> -- AVG = 715827881 (wrong) {code}
>  
>  
> Trigger 2: AVG over a join, where AVG is split into SUM/COUNT and pushed 
> below the join:
>  
> {code:java}
> from pyflink.table import EnvironmentSettings, TableEnvironment, 
> DataTypest_env = TableEnvironment.create(EnvironmentSettings.in_batch_mode())
> a = t_env.from_elements([(2147483647,), (2147483647,), (1,)], 
> DataTypes.ROW([DataTypes.FIELD('x', DataTypes.INT())]))
> b = t_env.from_elements([(1,), (2,)], DataTypes.ROW([DataTypes.FIELD('y', 
> DataTypes.INT())]))
> t_env.create_temporary_view('a', a)
> t_env.create_temporary_view('b', b)
> # Expected: 1431655765 = (2147483647 + 2147483647 + 1) / 3, each row repeated 
> twice by the join
> t_env.execute_sql("SELECT AVG(x) FROM a").print()                    # 
> 1431655765 (correct)
> t_env.execute_sql("SELECT AVG(x) FROM a, b").print()                 # 0      
>     (wrong)
> t_env.execute_sql("SELECT AVG(CAST(x AS BIGINT)) FROM a, b").print() # 
> 1431655765 (correct)
> print(t_env.explain_sql("SELECT AVG(x) FROM a, b"))
> {code}
>  



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to