[
https://issues.apache.org/jira/browse/SPARK-44027?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119678#comment-18119678
]
Alexander Fedorov commented on SPARK-44027:
-------------------------------------------
Hi everyone,
I'm a new contributor interested in this ticket, and I've spent some time
investigating the current state of the codebase. I wanted to share what I found
and propose an approach for feedback.
Current blockers I found:
SQLBuilder.scala was removed in SPARK-19025. There is no longer a
general-purpose SQL generator for logical plans in Spark.
The guards at views.scala:853 and DataSourceV2Strategy.scala:337 both require
originalText to be non-empty for persistent views. A DataFrame has no
originalText, so the feature is blocked architecturally, not just at the API
level.
LogicalPlan does not have a sql method; only small helper types like JoinType
and SortDirection have sql methods.
Proposed approach:
I'd like to explore a safe-subset implementation:
Generate SQL from a DataFrame's analyzed logical plan using a limited,
resurrected SQLBuilder.
Throw a clear AnalysisException for any plan that cannot be reliably
materialized (e.g., references to temporary views, session-scoped UDFs, or
unsupported operators).
Store the generated SQL as originalText, which satisfies the existing guards
without modifying them.
This is different from the original SQLBuilder, which attempted to handle every
operator. A safe-subset version would fail fast and explicitly rather than
producing SQL that breaks in a future session.
Questions:
Is a safe-subset approach acceptable, or does the community prefer a different
direction?
Are there specific operators or plan patterns that should be explicitly
excluded?
I'm happy to prototype this and share findings. Any guidance would be
appreciated.
> create *permanent* Spark View from DataFrame via PySpark & Scala DataFrame API
> ------------------------------------------------------------------------------
>
> Key: SPARK-44027
> URL: https://issues.apache.org/jira/browse/SPARK-44027
> Project: Spark
> Issue Type: New Feature
> Components: PySpark
> Affects Versions: 3.5.0
> Reporter: Martin Bode
> Priority: Major
> Labels: features, newbie
>
> currently only *_temporary_ Spark Views* can be created from a DataFrame:
> *
> [DataFrame.createTempView|https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.createTempView.html#pyspark.sql.DataFrame.createTempView]
> *
> [DataFrame.createOrReplaceTempView|https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.createOrReplaceTempView.html#pyspark.sql.DataFrame.createOrReplaceTempView]
> *
> [DataFrame.createGlobalTempView|https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.createGlobalTempView.html#pyspark.sql.DataFrame.createGlobalTempView]
> *
> [DataFrame.createOrReplaceGlobalTempView|https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.createOrReplaceGlobalTempView.html#pyspark.sql.DataFrame.createOrReplaceGlobalTempView]
> When a user needs a _*permanent*_ *Spark View* he has to fall back to Spark
> SQL ({{{}CREATE VIEW AS SELECT...{}}}).
> Sometimes it is easier and more readable to specify the desired logic of the
> view through {_}Scala/PySpark DataFrame API{_}.
> Therefore, I'd like to suggest to implement a new PySpark method that allows
> creating a _*permanent*_ *Spark View* from a DataFrame (e.g.
> {{{}DataFrame.createOrReplaceView{}}}).
> see also:
> *
> [https://community.databricks.com/s/question/0D53f00001PANVgCAP/is-there-a-way-to-create-a-nontemporary-spark-view-with-pyspark]
> * [https://lists.apache.org/thread/jzkznvt7cfjhmo77w1tlksxkwyvmvvfb]
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]