[
https://issues.apache.org/jira/browse/SPARK-58677?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Nathan Holland updated SPARK-58677:
-----------------------------------
Summary: Expand hash functionalities of XXH3 to better match XXHash64
(was: Expand hash functionalities of XXH3)
> Expand hash functionalities of XXH3 to better match XXHash64
> ------------------------------------------------------------
>
> Key: SPARK-58677
> URL: https://issues.apache.org/jira/browse/SPARK-58677
> Project: Spark
> Issue Type: Improvement
> Components: PySpark, SQL
> Affects Versions: 4.4.0
> Reporter: Nathan Holland
> Priority: Major
>
> This ticket expands on SPARK-45900, to explicitly mention the desire for a
> BINARY return type for xxhash3_128 as well as additional functionality that a
> modern hashing algorithm on Spark could implement. (such as deterministic
> support for complex types)
> The original intent of that ticket was to capture a new hashing model through
> XXH3 that allowed for high performance hashing using XXH3 with the API of the
> earlier XXHash64 hash family, rather than some of the cryptographic hash
> functions (SHA/MD5 etc)
> One difference in this model I am proposing - SPARK-45900 has been resolved
> by returning a string, which is more consistent with the older hashing
> models. When writing SPARK-45900 I mentioned that this was a downside of the
> existing SHA models, to clarifify - returning a STRING type has heavier
> overhead and prevents future optimisations that occur by allowing HotSpot
> (and C interpreters like Photon) to take advantage of SIMD instructions.
> Furthermore I've been experimenting with expanding the BINARY type to have
> fixed byte types and have seen some promising performance increases that may
> be useful at a later date.
> As XXH3 was built specifically around SIMD instructions to provide speed ups
> over XXHash64 and Spark 4.4 is not out yet, I believe now would be the time
> to discuss what this API should return - and either choose with a seperate
> API that returns the BINARY 128 bytes or make the change to the solution of
> SPARK-45900 to return a BINARY type by default and let users use the "hex"
> function if they want the result as a Binary String.
> I have heard from customers that one of xxhash64's benefits is the simple API
> to support multiple columns which conceptually allow the hashing of columns
> by passing them in natively without having to implement brittle solutions
> (often something like concat_ws) to hash data. I believe an expansion to the
> functionality to SPARK-45900 to support columns would allow XXH3 to be seen
> successor to the XXHash64 function with this user friendly API - while
> keeping all the semantics that a unitary function would support - including
> byte compatability with the original C reference code.
> My current codebase can be found here for a proposed implementation of the
> above two key changes (better support for hashing data types and multiple
> columns, and the return of a binary type which better captures the hash
> results rather than a string)
> *
> [https://github.com/NathanNZ/spark/commit/d4317b9853a9b1f775c1ab11534dc8ea5a048040]
> * I've also got a work in progress solution with pre-built containers to
> test the functionality.
> ** [https://github.com/NathanNZ/spark-nathan-containers]
> Links:
> - https://issues.apache.org/jira/browse/SPARK-45900 (Original SPARK-45900
> ticket)
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]