sunchao commented on PR #6427: URL: https://github.com/apache/datafusion-comet/pull/6427#issuecomment-5916719828
@viirya You’re right about Spark’s automatically injected runtime filters. I traced the original case, and my explanation was incorrect: **the literal comes from application code, not from Spark materializing a scalar subquery**. The application builds and serializes a Bloom filter, constructs `BloomFilterMightContain(Literal(bytes, BinaryType), key)`, and passes that expression through `DataFrame.filter`. That literal survives into `CometFilterExec.condition`. I reproduced this through an actual DataFrame query with native Comet execution, rather than manually constructing a physical filter. A 1,048,588-byte serialized Bloom produced a 2,097,226-character `SparkPlanInfo` string, while returning the expected results. The automatic runtime-filter controls retained ScalarSubquery with AQE both on and off. So the display issue is reachable, but the rationale should specifically describe application-built Bloom predicates. I’ll correct the description and add a real-query regression covering that path. I also agree that expression-level formatting in Spark would provide broader coverage, including fallback and projection cases. This patch is narrower: it summarizes these literals only in Comet filter displays. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
