hudi-agent commented on code in PR #20130:
URL: https://github.com/apache/hudi/pull/20130#discussion_r4139494541


##########
hudi-spark-datasource/hudi-spark/src/test/scala/org/apache/spark/sql/hudi/common/TestVectorizedReadWithSchemaEvolution.scala:
##########
@@ -17,53 +17,174 @@
 
 package org.apache.spark.sql.hudi.common
 
-import org.apache.hudi.HoodieSparkUtils
+import org.apache.spark.SparkException
 
 class TestVectorizedReadWithSchemaEvolution extends HoodieSparkSqlTestBase {
   Seq("cow", "mor").foreach { tableType =>
     test(s"Test vectorized read for $tableType table") {
-      if (HoodieSparkUtils.isSpark3) {
-        withSQLConf(
-          "hoodie.schema.on.read.enable" -> "true",
-          "spark.sql.parquet.enableVectorizedReader" -> "true",
-          "spark.sql.codegen.maxFields" -> "1",
-          "hoodie.parquet.small.file.limit" -> "0"
-        ) {
-          withTempDir { tmp =>
-            val tableName = generateTableName
-            val tablePath = s"${tmp.getCanonicalPath}/$tableName"
-            // create table
-            spark.sql(
-              s"""
-                 |create table $tableName (
-                 |  id int,
-                 |  name string,
-                 |  price double,
-                 |  ts long
-                 |) using hudi
-                 | partitioned by (ts)
-                 | location '$tablePath'
-                 | tblproperties (
-                 |  type = '$tableType',
-                 |  primaryKey = 'id',
-                 |  orderingFields = 'ts'
-                 | )
-         """.stripMargin)
-            // insert data to table
-            spark.sql(s"insert into $tableName values(1, 'a1', 10, 1000)")
-
-            // alter table change data type of column 'id'
-            spark.sql(s"alter table $tableName alter column price type string")
-
-            // insert new records
-            spark.sql(s"insert into $tableName values(2, 'a2', '20', 2000)")
-
-            checkAnswer(s"select id, name, price, ts from $tableName")(
-              Seq(1, "a1", "10.0", 1000),
-              Seq(2, "a2", "20", 2000)
-            )
-          }
+      withSQLConf(
+        "hoodie.schema.on.read.enable" -> "true",
+        "spark.sql.parquet.enableVectorizedReader" -> "true",
+        "spark.sql.codegen.maxFields" -> "1",
+        "hoodie.parquet.small.file.limit" -> "0"
+      ) {
+        withTempDir { tmp =>
+          val tableName = generateTableName
+          val tablePath = s"${tmp.getCanonicalPath}/$tableName"
+          // create table
+          spark.sql(
+            s"""
+                |create table $tableName (
+                |  id int,
+                |  name string,
+                |  price double,
+                |  ts long
+                |) using hudi
+                | partitioned by (ts)
+                | location '$tablePath'
+                | tblproperties (
+                |  type = '$tableType',
+                |  primaryKey = 'id',
+                |  orderingFields = 'ts'
+                | )
+        """.stripMargin)
+          // insert data to table
+          spark.sql(s"insert into $tableName values(1, 'a1', 10, 1000)")
+
+          // alter table change data type of column 'id'
+          spark.sql(s"alter table $tableName alter column price type string")
+
+          // insert new records
+          spark.sql(s"insert into $tableName values(2, 'a2', '20', 2000)")
+
+          checkAnswer(s"select id, name, price, ts from $tableName")(
+            Seq(1, "a1", "10.0", 1000),
+            Seq(2, "a2", "20", 2000)
+          )
+        }
+      }
+    }
+  }
+
+  test(s"Test INT to DECIMAL schema evolution with precision overflow for COW 
table") {

Review Comment:
   🤖 nit: the three new overflow tests repeat the same `withSQLConf` block and 
`create table` DDL. Could you pull those into a small helper that takes the 
source column type and ANSI flag?
   
   <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag 
quality.</i></sub>



##########
hudi-client/hudi-spark-client/src/main/java/org/apache/hudi/client/utils/SparkInternalSchemaConverter.java:
##########
@@ -325,6 +326,22 @@ private static DataType constructSparkSchemaFromType(Type 
type) {
     }
   }
 
+  private static void putDecimalOrNull(

Review Comment:
   🤖 nit: `putDecimalOrNull` doesn't mention that this throws when ANSI mode is 
on. Could you rename it (e.g. `putDecimalWithOverflowHandling`) or add a short 
Javadoc that says it nulls the value in non-ANSI mode and throws in ANSI mode?
   
   <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag 
quality.</i></sub>



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to