[
https://issues.apache.org/jira/browse/SPARK-59717?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59717:
-----------------------------------
Labels: pull-request-available (was: )
> GROUP BY / DISTINCT on a UDT backed by a primitive type fails with
> scala.MatchError
> -----------------------------------------------------------------------------------
>
> Key: SPARK-59717
> URL: https://issues.apache.org/jira/browse/SPARK-59717
> Project: Spark
> Issue Type: Bug
> Components: SQL
> Affects Versions: 5.0.0
> Reporter: Eric Yang
> Priority: Major
> Labels: pull-request-available
>
> Grouping by, or calling distinct() on a column of a UserDefinedType whose
> sqlType is a primitive type (e.g. IntegerType, DoubleType) fails with default
> settings:
> {code:java}
> scala.MatchError: org.apache.spark.sql.YearUDT@... (of class
> org.apache.spark.sql.YearUDT)
> at
> org.apache.spark.sql.execution.aggregate.HashMapGenerator.genComputeHash{code}
> *Repro* (with YearUDT from UserDefinedTypeSuite, sqlType = IntegerType):
> {code:java}
> UDTRegistration.register(classOf[java.time.Year].getName,
> classOf[YearUDT].getName)
> Seq((Year.of(2024), 1), (Year.of(2025), 2)).toDF("y",
> "v").groupBy("y").count().collect()
> Through collect() the error is reported as [INTERNAL_ERROR]. The same query
> works with spark.sql.codegen.wholeStage=false.{code}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]