Hi all, I'd like to get feedback from the dev list on SPARK-58126 / PR #57257: https://github.com/apache/spark/pull/57257
Motivation: Pool/TaskSet ordering is currently hard-wired to FairSchedulingAlgorithm or FIFOSchedulingAlgorithm, selected only via spark.scheduler.mode. Users who need custom ordering (priority-based, deadline-aware, SLA-driven, ...) have no supported extension point today and must patch Spark. This is especially relevant on managed platforms (Databricks, EMR, ...) where users don't create the SparkContext themselves, so a runtime setter isn't an option. Proposed change: A new optional config, spark.scheduler.rootPool.comparator.class, lets users plug in a java.util.Comparator[SchedulableInfo] to order the root pool. SchedulableInfo is a new, immutable @DeveloperApi snapshot type (scheduling mode, weight, min share, running tasks, priority, stage id, name) — the only type exposed to user code. Spark's internal, mutable Schedulable/SchedulingAlgorithm hierarchy stays private[spark] and free to evolve. When the config is unset, behavior is unchanged. This follows the existing pluggable-class pattern used by spark.serializer and spark.shuffle.manager. Open question: Is @DeveloperApi the right stability level for this extension point, or would the community prefer this go through a SPIP given it introduces a new public-facing type and config? I'm happy to go either route and would appreciate input from anyone familiar with the scheduler. Tests: New PoolSuite coverage for direct comparator injection and config-based root-pool override, in both FAIR and FIFO mode (16/16 passing). Thanks for taking a look. Best, Christian --------------------------------------------------------------------- To unsubscribe e-mail: [email protected]
