andygrove opened a new issue, #5929:
URL: https://github.com/apache/datafusion-comet/issues/5929

   ### What is the problem the feature request solves?
   
   `ci.yml` runs `pr_build_linux` on push to main. That is deliberate and
   documented: an `actions/cache` entry can only be restored from the current
   branch or from the default branch, and the merge queue builds on a throwaway
   `gh-readonly-queue/*` branch whose caches are deleted with it. Without a push
   run, main's caches go stale and every later pull request pays the delta.
   
   The problem is that it refreshes those caches by running the *entire*
   pipeline. Averaged over four recent push-to-main runs, `pr_build_linux` costs
   587 runner-minutes a push, and all but about 73 of that is lints, the 5x4
   `linux-test` matrix and the TPC-H/TPC-DS query passes. None of it tests
   anything new: the merge queue already ran the same jobs against the exact 
tree
   that landed, usually less than an hour earlier.
   
   | job group                             | avg min/push | writes a cache main 
needs?        |
   | ------------------------------------- | -----------: | 
--------------------------------- |
   | `linux-test` (5 profiles x 4 suites)  |        456.1 | no                  
              |
   | `Build Native Library`                |         32.3 | yes, `cargo-ci`     
              |
   | `Verify TPC-DS Results`               |         26.4 | yes, dataset + 
`java-maven`       |
   | `ubuntu-latest/rust-test`             |         23.7 | yes, `cargo-debug`  
              |
   | `Lint Java` (matrix)                  |         22.0 | no, prefix-shares 
`java-maven`    |
   | `Celeborn` (2 versions)               |         10.8 | no, prefix-shares 
`java-maven`    |
   | `Verify TPC-H Results`                |         10.4 | yes, dataset + 
`java-maven`       |
   | `Build Spark 4.1, JDK 17`             |          4.5 | no, prefix-shares 
`java-maven`    |
   | `Lint`                                |          0.7 | no, but the two 
native jobs need it |
   | `Lint Scala (syntactic)`              |          0.5 | no                  
              |
   | **total**                             |    **587.4** |                     
              |
   
   Inside the two `Verify` jobs, the query passes are themselves most of the
   runtime and none of it feeds a cache: 18.1 of TPC-DS's 26.4 minutes and 2.6 
of
   TPC-H's 10.4.
   
   At the 8-10 pushes a day this repository sees, that is roughly 4,100-5,200
   runner-minutes a day, or 70-85 runner-hours, spent re-testing what the queue
   just tested. For scale, all of `ci.yml` came to 1,257 runner-hours on
   2026-09-12 (44,440 minutes on pull requests, 26,395 in the queue, 4,588 on
   push).
   
   ### Describe the potential solution
   
   Split the push tier from the test tier. Keep `pr_build_linux` on `push` so 
the
   caches stay warm, but run only the jobs that write one, and each only far
   enough to populate it:
   
   - `build-native` — `cargo-ci`
   - `linux-test-rust` — `cargo-debug`. The compile is nearly all of its 
runtime,
     so a compile-only variant would save almost nothing and would risk the 
cache
     contents diverging from what a pull request restores.
   - `verify-benchmark-results-tpch` / `-tpcds` — the SF=1 datasets, expensive 
to
     regenerate, plus the shared `java-maven` entry. Skip only the query passes;
     data generation still has to run.
   - `lint` — 40 seconds, and both native jobs `needs:` it.
   
   Everything else gets `if: ${{ !inputs.cache-refresh-only }}`. That takes the
   push tier from 587 runner-minutes to about 73.
   
   Mechanically this is the existing `spark_4_1` / `spark_4_1_hive` pattern: a
   second POLICY output (`build_linux_full`, tiers `["pr", "queue"]`) feeding 
the
   same call, folded into a `cache-refresh-only` input on the reusable workflow.
   Using an input rather than a separate cache-warming workflow keeps every 
cache
   key in one file, where two copies cannot drift apart.
   
   `dev/ci/check-ci-config.py` should fail if a job is added to
   `pr_build_linux.yml` without either the guard or an entry naming the cache it
   writes. Otherwise the push tier quietly grows back and nothing goes red.
   
   ### Additional context
   
   Two deliberate consequences worth stating:
   
   - The suffixed Maven keys (`-lint`, `-spark-4.1-build`, `-celeborn-*`) stop
     being refreshed on main and fall back to the shared `<os>-java-maven-`
     restore-keys prefix, which the TPC jobs still write. The cost is a small
     delta download per lint job on a pull request.
   - `<os>-cargo-registry-*` stops being written on main, because `linux-test`
     owns it. Those jobs pass `skip-native-build: true` and run no cargo 
command,
     so the entry is close to unused and a miss changes nothing.
   
   This is the first of several CI cost items. Trimming redundant Spark profiles
   from the PR-tier `linux-test` matrix, and moving Iceberg 1.11 behind the
   existing `run-iceberg-tests` label, are larger but need their own evidence 
and
   should be argued separately.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to