Venkata krishnan Sowrirajan created SPARK-59806:
---------------------------------------------------
Summary: Make shuffle tracking storage-aware for reliably-stored
shuffles
Key: SPARK-59806
URL: https://issues.apache.org/jira/browse/SPARK-59806
Project: Spark
Issue Type: Improvement
Components: Spark Core
Affects Versions: 4.2.0
Reporter: Venkata krishnan Sowrirajan
When dynamic allocation's shuffle tracking
(spark.dynamicAllocation.shuffleTracking.enabled) is on together with a
reliable remote shuffle service, Spark tracks every shuffle-producing executor
without consulting where the shuffle output actually lives. Executors that
wrote reliably-stored output stay pinned for
spark.dynamicAllocation.shuffleTracking.timeout even though that output
survives executor removal, so they are never scaled down.
This is worst in mixed/fallback setups: local-fallback shuffles genuinely need
tracking, remote ones do not, and the app-global reliable-storage flag cannot
tell them apart.
Building on the per-shuffle reliability state from SPARK-59138, make
ExecutorMonitor consult MapOutputTracker.isReliablyStored(shuffleId) when
handling SparkListenerJobStart, and exclude reliably-stored shuffles from
shuffle tracking while retaining tracking for local or unknown-storage shuffles.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]