andygrove opened a new issue, #6386:
URL: https://github.com/apache/datafusion-comet/issues/6386
### What is the problem the feature request solves?
The merge queue builds at most two entries at once (`max_entries_to_build:
2` in `.asf.yaml`). Between 2026-09-11 and 2026-09-29 a full queue build took a
median of 92 minutes (p90 117), so even a queue that never runs dry tops out at
roughly 30 merges a day. Over the last week PRs were approved at about 20 a day
and merged at about 13.6 a day, so the backlog keeps growing, and when a batch
of PRs is queued together the last one lands 10 to 12 hours later.
A simulation that replays the observed build durations and the observed ~12%
per-entry failure rate gives:
| `max_entries_to_build` | Merges/day, queue never empty | Build time lost
to rebuilds after a failure |
| --- | --- | --- |
| 2 | ~30 | ~11% |
| 3 | ~44 | ~14% |
| 4 | ~58 | ~16% |
| 6 | ~82 | ~21% |
### Describe the potential solution
Set `max_entries_to_build: 4`. The comment on the setting already
anticipates this ("Raise it if the queue backs up"). The runner cost of a merge
does not change. What grows is the waste when an entry fails, because the
entries behind it that were already building have to start again.
While there, correct the comment on the `merge_group` branch of the `Detect
changes` step in `ci.yml`. It says `merge_group.base_sha` covers every entry
batched into the group, but `base_sha` is the queue commit of the entry ahead,
not `main`. For example, `gh-readonly-queue/main/pr-6357-1e08e6c9` was built on
`1e08e6c9`, which was #6300's queue commit. Each entry's build therefore runs
only the suites that its own changed files select. That is sound under
`ALLGREEN`, and it is why switching to `HEADGREEN` would not be: a green entry
could carry an earlier entry whose suites never ran.
### Additional context
Runner budget: an estimated ~190k runner-minutes a week today (PR runs ~96k,
queue ~68k, label runs ~10k, nightly ~9k) against the ASF cap of 250k a week.
Concurrency does not change the cost per merge, but merging more PRs a week
does.
Related: #5969 (cancel a queue run on its first failure, which frees the
slot sooner).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]