[
https://issues.apache.org/jira/browse/FLINK-40919?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Yaoxuan Wu updated FLINK-40919:
-------------------------------
Description:
{{AVG}} on an INT column is computed without overflow on its own. When the
planner rewrites it into {{{}SUM / COUNT{}}}, the sum keeps type INT and wraps
around, so the same AVG over the same rows returns a wrong value.
Trigger 1: SUM and COUNT of the same column in the SELECT list:
{code:java}
SELECT AVG(x) FROM (VALUES (2147483647), (2147483647), (2147483647)) AS v(x);
-- 2147483647 (correct)
SELECT AVG(x), SUM(x), COUNT(x) FROM (VALUES (2147483647), (2147483647),
(2147483647)) AS v(x);
-- AVG = 715827881 (wrong) {code}
Trigger 2: AVG over a join, where AVG is split into SUM/COUNT and pushed below
the join:
{code:java}
from pyflink.table import EnvironmentSettings, TableEnvironment, DataTypest_env
= TableEnvironment.create(EnvironmentSettings.in_batch_mode())
a = t_env.from_elements([(2147483647,), (2147483647,), (1,)],
DataTypes.ROW([DataTypes.FIELD('x', DataTypes.INT())]))
b = t_env.from_elements([(1,), (2,)], DataTypes.ROW([DataTypes.FIELD('y',
DataTypes.INT())]))
t_env.create_temporary_view('a', a)
t_env.create_temporary_view('b', b)
# Expected: 1431655765 = (2147483647 + 2147483647 + 1) / 3, each row repeated
twice by the join
t_env.execute_sql("SELECT AVG(x) FROM a").print() #
1431655765 (correct)
t_env.execute_sql("SELECT AVG(x) FROM a, b").print() # 0
(wrong)
t_env.execute_sql("SELECT AVG(CAST(x AS BIGINT)) FROM a, b").print() #
1431655765 (correct)
print(t_env.explain_sql("SELECT AVG(x) FROM a, b"))
{code}
was:
{{AVG(x)}} on an INT column is computed without overflow on its own. When the
planner rewrites it into {{{}SUM(x) / COUNT(x){}}}, the sum keeps type INT and
wraps around, so the same AVG over the same rows returns a wrong value.
Trigger 1: SUM and COUNT of the same column in the SELECT list:
{code:java}
SELECT AVG(x) FROM (VALUES (2147483647), (2147483647), (2147483647)) AS v(x);
-- 2147483647 (correct)
SELECT AVG(x), SUM(x), COUNT(x) FROM (VALUES (2147483647), (2147483647),
(2147483647)) AS v(x);
-- AVG = 715827881 (wrong) {code}
Trigger 2: AVG over a join, where AVG is split into SUM/COUNT and pushed below
the join:
{code:java}
from pyflink.table import EnvironmentSettings, TableEnvironment, DataTypest_env
= TableEnvironment.create(EnvironmentSettings.in_batch_mode())
a = t_env.from_elements([(2147483647,), (2147483647,), (1,)],
DataTypes.ROW([DataTypes.FIELD('x', DataTypes.INT())]))
b = t_env.from_elements([(1,), (2,)], DataTypes.ROW([DataTypes.FIELD('y',
DataTypes.INT())]))
t_env.create_temporary_view('a', a)
t_env.create_temporary_view('b', b)
# Expected: 1431655765 = (2147483647 + 2147483647 + 1) / 3, each row repeated
twice by the join
t_env.execute_sql("SELECT AVG(x) FROM a").print() #
1431655765 (correct)
t_env.execute_sql("SELECT AVG(x) FROM a, b").print() # 0
(wrong)
t_env.execute_sql("SELECT AVG(CAST(x AS BIGINT)) FROM a, b").print() #
1431655765 (correct)
print(t_env.explain_sql("SELECT AVG(x) FROM a, b"))
{code}
> AVG on INT returns wrong result when the planner rewrites it into SUM/COUNT
> (INT sum overflows)
> -----------------------------------------------------------------------------------------------
>
> Key: FLINK-40919
> URL: https://issues.apache.org/jira/browse/FLINK-40919
> Project: Flink
> Issue Type: Bug
> Components: Table SQL / Planner
> Affects Versions: 2.3.0
> Reporter: Yaoxuan Wu
> Priority: Major
>
> {{AVG}} on an INT column is computed without overflow on its own. When the
> planner rewrites it into {{{}SUM / COUNT{}}}, the sum keeps type INT and
> wraps around, so the same AVG over the same rows returns a wrong value.
> Trigger 1: SUM and COUNT of the same column in the SELECT list:
>
> {code:java}
> SELECT AVG(x) FROM (VALUES (2147483647), (2147483647), (2147483647)) AS v(x);
> -- 2147483647 (correct)
> SELECT AVG(x), SUM(x), COUNT(x) FROM (VALUES (2147483647), (2147483647),
> (2147483647)) AS v(x);
> -- AVG = 715827881 (wrong) {code}
>
>
> Trigger 2: AVG over a join, where AVG is split into SUM/COUNT and pushed
> below the join:
>
> {code:java}
> from pyflink.table import EnvironmentSettings, TableEnvironment,
> DataTypest_env = TableEnvironment.create(EnvironmentSettings.in_batch_mode())
> a = t_env.from_elements([(2147483647,), (2147483647,), (1,)],
> DataTypes.ROW([DataTypes.FIELD('x', DataTypes.INT())]))
> b = t_env.from_elements([(1,), (2,)], DataTypes.ROW([DataTypes.FIELD('y',
> DataTypes.INT())]))
> t_env.create_temporary_view('a', a)
> t_env.create_temporary_view('b', b)
> # Expected: 1431655765 = (2147483647 + 2147483647 + 1) / 3, each row repeated
> twice by the join
> t_env.execute_sql("SELECT AVG(x) FROM a").print() #
> 1431655765 (correct)
> t_env.execute_sql("SELECT AVG(x) FROM a, b").print() # 0
> (wrong)
> t_env.execute_sql("SELECT AVG(CAST(x AS BIGINT)) FROM a, b").print() #
> 1431655765 (correct)
> print(t_env.explain_sql("SELECT AVG(x) FROM a, b"))
> {code}
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)