[
https://issues.apache.org/jira/browse/SPARK-59546?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59546:
-----------------------------------
Labels: pull-request-available (was: )
> Avoid per-row exceptions when inferring TimeType in JSON, CSV, and XM
> ---------------------------------------------------------------------
>
> Key: SPARK-59546
> URL: https://issues.apache.org/jira/browse/SPARK-59546
> Project: Spark
> Issue Type: Improvement
> Components: SQL
> Affects Versions: 4.1.0
> Reporter: David Mollitor
> Priority: Minor
> Labels: pull-request-available
>
> h2. Summary
> JSON, CSV, and XML schema inference probe whether a string value is a
> {{TIME}} by calling the throwing {{TimeFormatter.parse}} inside {{allCatch
> opt}}:
> {code:scala}
> if (isTimeTypeEnabled && (allCatch opt timeFormatter.parse(field)).isDefined)
> {
> TimeType(TimeType.DEFAULT_PRECISION)
> }
> {code}
> For the default (pattern-less) formatter, {{parse}} calls
> {{SparkDateTimeUtils.stringToTimeAnsi}}, which throws a
> {{SparkDateTimeException}} (filling in a stack trace) for every value that is
> not a time. During inference of a non-time column, this is one
> thrown-and-caught exception per row.
> This is the same cost that SPARK-39193 / SPARK-39280 / SPARK-39281 removed
> for timestamp and date inference by introducing a non-throwing
> {{parseOptional}} (that change reported exception handling as more than 90%
> of type-inference time). {{TimeType}} was added later and never got a
> {{parseOptional}}, so it is the only datetime formatter still using the
> throwing probe -- the timestamp probe immediately below the time probe
> already uses {{parseOptional}}.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]