[
https://issues.apache.org/jira/browse/SPARK-58275?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-58275:
-----------------------------------
Labels: pull-request-available (was: )
> Add normalize Unicode normalization SQL function
> ------------------------------------------------
>
> Key: SPARK-58275
> URL: https://issues.apache.org/jira/browse/SPARK-58275
> Project: Spark
> Issue Type: New Feature
> Components: SQL
> Affects Versions: 5.0.0
> Reporter: sreeramayeshwanthgowd
> Priority: Major
> Labels: pull-request-available
>
> Spark has no built in function for Unicode normalization. Today users have to
> fall back to a UDF, which cannot be optimized by Catalyst and tends to be
> slower. It would be helpful to expose Unicode normalization directly in SQL
> and the DataFrame APIs.
> This proposes adding a built in normalize(str[, form]) scalar function that
> returns the Unicode normalization of the input string using the given
> normalization form. Supported forms are NFC (the default), NFD, NFKC, and
> NFKD, and the form name is case insensitive. It can be implemented with
> java.text.Normalizer, so no new dependency is required.
> normalize is the ANSI SQL standard NORMALIZE function and is already
> available in PostgreSQL, Trino, and Google BigQuery, so this also improves
> parity with the engines that Spark users migrate from.
> Proposed interface:
> SQL: normalize(str) and normalize(str, form)
> Scala: functions.normalize(col) and functions.normalize(col, form)
> PySpark (classic and Spark Connect): pyspark.sql.functions.normalize(str,
> form=None)
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]