andygrove opened a new issue, #6388:
URL: https://github.com/apache/datafusion-comet/issues/6388

   ### What is the problem the feature request solves?
   
   The Spark SQL suite sets the length of every merge-queue run: its build job 
(18 minutes with a warm cache) plus its slowest shard. Median shard durations 
across 129 queue runs between 2026-09-11 and 2026-09-29:
   
   | Row | Median |
   | --- | --- |
   | `sql_hive-1` | 66.7 min |
   | `sql_core-1` | 59.6 min |
   | `sql_core-2` | 48.2 min |
   | `sql_hive-3` | 44.9 min |
   | `sql_core-3` | 28.1 min |
   | `sql_hive-2` | 21.1 min |
   | `catalyst` | 14.0 min |
   
   Two things make the long rows long:
   
   - `org.apache.spark.sql.hive.client.HivePartitionFilteringSuites` runs 
`HivePartitionFilteringSuite` against every Hive client version and took 21 to 
35 minutes of `sql_hive-1` in the runs I sampled. It is a single class, so a 
name filter cannot split it further.
   - `org.apache.spark.sql.execution.datasources.*` and 
`org.apache.spark.sql.connector.*` account for 13 to 19 minutes of 
`sql_core-1`, and the data source suites for another 11 minutes of `sql_core-2`.
   
   ### Describe the potential solution
   
   Move them into two new rows with sbt's exclusion patterns (`testOnly * 
-<glob>`), keeping the tag filters so every test still runs exactly once:
   
   - `sql_core-4`: `sql/testOnly org.apache.spark.sql.execution.datasources.* 
org.apache.spark.sql.connector.* -- -l org.apache.spark.tags.SlowSQLTest` 
(their untagged and Extended tests), with both globs excluded from `sql_core-1` 
and `sql_core-2`.
   - `sql_hive-4`: `hive/testOnly 
org.apache.spark.sql.hive.client.HivePartitionFilteringSuites` with 
`sql_hive-1`'s tag filters, and the class excluded from `sql_hive-1`.
   
   In the two runs I sampled, the slowest row would drop from 72 and 51 minutes 
to about 48 and 40. The cost is two more runners per Spark SQL run, each with 
about 9 minutes of setup.
   
   ### Additional context
   
   #6103 / #6113 shard the Iceberg extensions task, which is the next longest 
path (about 57 minutes).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to