Hi all,

I'd like to get feedback from the dev list on SPARK-58126 / PR #57257:
https://github.com/apache/spark/pull/57257

Motivation:
Pool/TaskSet ordering is currently hard-wired to FairSchedulingAlgorithm or 
FIFOSchedulingAlgorithm, selected only via spark.scheduler.mode. Users who need 
custom ordering (priority-based, deadline-aware, SLA-driven, ...) have no 
supported extension point today and must patch Spark. This is especially 
relevant on managed platforms (Databricks, EMR, ...) where users don't create 
the SparkContext themselves, so a runtime setter isn't an option.

Proposed change:
A new optional config, spark.scheduler.rootPool.comparator.class, lets users 
plug in a java.util.Comparator[SchedulableInfo] to order the root pool. 
SchedulableInfo is a new, immutable @DeveloperApi snapshot type (scheduling 
mode, weight, min share, running tasks, priority, stage id, name) — the only 
type exposed to user code. Spark's internal, mutable 
Schedulable/SchedulingAlgorithm hierarchy stays private[spark] and free to 
evolve. When the config is unset, behavior is unchanged.

This follows the existing pluggable-class pattern used by spark.serializer and 
spark.shuffle.manager.

Open question:
Is @DeveloperApi the right stability level for this extension point, or would 
the community prefer this go through a SPIP given it introduces a new 
public-facing type and config? I'm happy to go either route and would 
appreciate input from anyone familiar with the scheduler.

Tests: New PoolSuite coverage for direct comparator injection and config-based 
root-pool override, in both FAIR and FIFO mode (16/16 passing).

Thanks for taking a look.

Best,
Christian

---------------------------------------------------------------------
To unsubscribe e-mail: [email protected]

Reply via email to