Venkata krishnan Sowrirajan created SPARK-59806:
---------------------------------------------------

             Summary:  Make shuffle tracking storage-aware for reliably-stored 
shuffles
                 Key: SPARK-59806
                 URL: https://issues.apache.org/jira/browse/SPARK-59806
             Project: Spark
          Issue Type: Improvement
          Components: Spark Core
    Affects Versions: 4.2.0
            Reporter: Venkata krishnan Sowrirajan


When dynamic allocation's shuffle tracking 
(spark.dynamicAllocation.shuffleTracking.enabled) is on together with a 
reliable remote shuffle service, Spark tracks every shuffle-producing executor 
without consulting where the shuffle output actually lives. Executors that 
wrote reliably-stored output stay pinned for 
spark.dynamicAllocation.shuffleTracking.timeout even though that output 
survives executor removal, so they are never scaled down.

This is worst in mixed/fallback setups: local-fallback shuffles genuinely need 
tracking, remote ones do not, and the app-global reliable-storage flag cannot 
tell them apart.

Building on the per-shuffle reliability state from SPARK-59138, make 
ExecutorMonitor consult MapOutputTracker.isReliablyStored(shuffleId) when 
handling SparkListenerJobStart, and exclude reliably-stored shuffles from 
shuffle tracking while retaining tracking for local or unknown-storage shuffles.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to