[
https://issues.apache.org/jira/browse/SPARK-58857?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-58857:
-----------------------------------
Labels: pull-request-available (was: )
> LocalDirsFeatureStep discards the result of Utils.randomize, so configured
> local dirs are never shuffled
> --------------------------------------------------------------------------------------------------------
>
> Key: SPARK-58857
> URL: https://issues.apache.org/jira/browse/SPARK-58857
> Project: Spark
> Issue Type: Bug
> Components: Kubernetes
> Affects Versions: 5.0.0
> Reporter: Yang Jie
> Priority: Major
> Labels: pull-request-available
>
> `LocalDirsFeatureStep.configurePod` calls `Utils.randomize` on the local dirs
> resolved from configuration but discards the result, so that branch never
> randomizes.
> {code}
> val resolvedLocalDirs = Option(conf.sparkConf.getenv("SPARK_LOCAL_DIRS"))
> .orElse(conf.getOption("spark.local.dir"))
> .getOrElse(defaultLocalDir)
> .split(",")
> randomize(resolvedLocalDirs) // result dropped
> localDirs = resolvedLocalDirs.toImmutableArraySeq // original order
> {code}
> `Utils.randomize[T](seq: IterableOnce[T]): Seq[T]` is
> `randomizeInPlace(seq.iterator.toArray).toImmutableArraySeq` - it shuffles a
> fresh copy and returns it, so the caller observes no change to the argument.
> The in-place variant is `Utils.randomizeInPlace`. The sibling call fifteen
> lines earlier, for the pod-template branch, binds the result correctly:
> {code}
> var localDirs = randomize(pod.container.getVolumeMounts.asScala
> .filter(_.getName.startsWith("spark-local-dir-"))
> .map(_.getMountPath))
> {code}
> Both calls arrived in the same commit, 9f9af2a7bf6 (SPARK-39755, "Improve
> LocalDirsFeatureStep to randomize local directories"). An earlier revision of
> that PR used `randomizeInPlace`; review suggested `Utils.randomize` instead,
> and the next revision adopted it at both sites, binding the result at one and
> dropping it at the other. During that review a committer twice noted the test
> did not fail as expected, which is the symptom of this bug, but the
> discrepancy was attributed to force-pushes and the assertions were loosened
> to accept either order.
> So `SPARK_LOCAL_DIRS`, the emptyDir volume list, and the `spark-local-dir-N`
> to path pairing are emitted in configured order, identically for every driver
> and executor pod in an application.
> Scope of the observable effect, so this is not oversold: no measurable I/O or
> capacity skew follows. `DiskBlockManager.getFile` picks a directory by
> `nonNegativeHash(filename) % localDirs.length`, which spreads files
> near-uniformly within each executor regardless of the order it receives, and
> in this branch every path is backed by an emptyDir volume this step creates,
> so all of them share one node filesystem (or RAM under
> `spark.kubernetes.local.dirs.tmpfs`). What is real is that the randomization
> SPARK-39755 added, and its JIRA and release note describe, does not happen on
> this branch, and the suite's `// SPARK-39755 : Changes the method to test
> randomization` comment describes behavior that never occurs.
> Fix: bind the result, matching the sibling call. Deleting the call as dead
> code is a defensible alternative if the randomization is not considered worth
> having.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]