[
https://issues.apache.org/jira/browse/SPARK-59124?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
WonYong-Jang updated SPARK-59124:
---------------------------------
Description:
In K8s cluster mode, {{SparkSubmit}} unconditionally downloads {{spark.files}}
to the driver's
local disk and rewrites the URIs to
{{{}[file:/tmp/spark-|file:///tmp/spark-]<uuid>/...{}}}, discarding the original
remote URI (e.g. the {{s3a://}} path under
{{{}spark.kubernetes.file.upload.path{}}}). The files
are still delivered to executors via the driver file server, so
{{SparkFiles.get()}} works
inside a task, but no single shared path remains that {{spark.read.*}} can use.
SPARK-47475 already made this skippable for jars via
{{spark.kubernetes.jars.avoidDownloadSchemes}} (4.0.0), and SPARK-55077 is
doing the same for
archives. This asks for the same for files, via
{{{}spark.kubernetes.files.avoidDownloadSchemes{}}}, applying the existing
{{avoidDownload}}
predicate to {{args.files}} in {{{}SparkSubmit.prepareSubmitEnvironment{}}}.
Full description, code references, reproduction and workaround:
[https://github.com/apache/spark/issues/58310]
was:
h2. Summary
In K8s cluster mode, {{SparkSubmit}} unconditionally downloads {{spark.files}}
to the driver's
local disk and rewrites the URIs to
{{{}[file:/tmp/spark-|file:///tmp/spark-]<uuid>/...{}}}, discarding the original
remote URI (e.g. the {{s3a://}} path under
{{{}spark.kubernetes.file.upload.path{}}}). The files
are still delivered to executors via the driver file server, so
{{SparkFiles.get()}} works
inside a task, but no single shared path remains that {{spark.read.*}} can use.
SPARK-47475 already made this skippable for jars via
{{spark.kubernetes.jars.avoidDownloadSchemes}} (4.0.0), and SPARK-55077 is
doing the same for
archives. This asks for the same for files, via
{{{}spark.kubernetes.files.avoidDownloadSchemes{}}}, applying the existing
{{avoidDownload}}
predicate to {{args.files}} in {{{}SparkSubmit.prepareSubmitEnvironment{}}}.
Full description, code references, reproduction and workaround:
[https://github.com/apache/spark/issues/58310]
> Support spark.kubernetes.files.avoidDownloadSchemes for K8s cluster mode
> ------------------------------------------------------------------------
>
> Key: SPARK-59124
> URL: https://issues.apache.org/jira/browse/SPARK-59124
> Project: Spark
> Issue Type: Improvement
> Components: Kubernetes, Spark Core
> Affects Versions: 5.0.0
> Reporter: WonYong-Jang
> Priority: Major
>
> In K8s cluster mode, {{SparkSubmit}} unconditionally downloads
> {{spark.files}} to the driver's
> local disk and rewrites the URIs to
> {{{}[file:/tmp/spark-|file:///tmp/spark-]<uuid>/...{}}}, discarding the
> original
> remote URI (e.g. the {{s3a://}} path under
> {{{}spark.kubernetes.file.upload.path{}}}). The files
> are still delivered to executors via the driver file server, so
> {{SparkFiles.get()}} works
> inside a task, but no single shared path remains that {{spark.read.*}} can
> use.
> SPARK-47475 already made this skippable for jars via
> {{spark.kubernetes.jars.avoidDownloadSchemes}} (4.0.0), and SPARK-55077 is
> doing the same for
> archives. This asks for the same for files, via
> {{{}spark.kubernetes.files.avoidDownloadSchemes{}}}, applying the existing
> {{avoidDownload}}
> predicate to {{args.files}} in {{{}SparkSubmit.prepareSubmitEnvironment{}}}.
> Full description, code references, reproduction and workaround:
> [https://github.com/apache/spark/issues/58310]
>
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]