goutamadwant opened a new issue, #12482:
URL: https://github.com/apache/seatunnel/issues/12482
### Search before asking
I searched open issues and PRs:
- #11513 (umbrella) keeps "Linux JDK 8/11 as the normal PR compatibility
matrix" for unit tests. This proposal is only about the connector E2E jobs.
- #11545 moves the CI JDKs to 11/17. That changes which two JDKs run, not
how many E2E legs a PR starts; the proposal applies either way.
- #11515 and #9976 do not change the JDK matrix.
### Summary
Every connector E2E job runs twice, on JDK 8 and JDK 11 (`matrix.java: ['8',
'11']`). The host JDK runs Maven, the Testcontainers client and the test
assertions. The engines under test run inside containers with their own JVM:
the Zeta container is `seatunnelhub/openjdk:8u342`
(`SeaTunnelContainer.java:95`), and Flink/Spark use their images' JVMs. So the
second leg mostly repeats the first against the same engine JVMs.
### Evidence
Scope: 66 full-matrix `Build` runs on one contributor fork, the latest 100
completed runs as of 2026-09-24, restricted to runs where the connector shards
ran. Read-only `gh api` job data. Intervals over 320 minutes were dropped as
invalid.
| Job family | JDK 8 runner-min / run | JDK 11 runner-min / run | JDK pairs
| Failed only on JDK 8 | Failed only on JDK 11 |
|---|--:|--:|--:|--:|--:|
| Connector E2E (all-connectors, JDBC, dedicated, updated-modules) | 2,545 |
2,526 | 1,813 | 27 | 16 |
| Transform E2E | 231 | 230 | 121 | 8 | 9 |
| Engine E2E | 144 | 133 | 111 | 12 | 12 |
- The connector E2E JDK 11 leg is about 2,500 runner-minutes per full run,
roughly half of all IT time (median 5,169 IT runner-min per run).
- I read the logs of the 16 connector failures that happened only on JDK 11:
- 11 match known flaky or shared signatures that also fail on JDK 8
elsewhere: `PostgresCDCIT` committed-offset ×4, `PaimonWithS3IT` privilege ×2,
Doris condition timeouts ×2, `NebulaGraphIT` startup, `DatabendCDCSinkIT`, and
the OceanBase CDC `TableId` regression.
- 1 `RocketMqIT` assertion needs a closer look.
- 1 was `JdbcPhoenixIT` class setup.
- 3 had no test-failure line (build or infrastructure).
- None showed a JDK-specific error such as a class-version, module-access
or API difference.
- One-JDK-only failures are balanced across the two JDKs (27 vs 16), which
points to flakiness rather than JDK-specific behaviour.
### Proposal
1. In PR runs, run connector E2E jobs (the all-connectors shards, JDBC
parts, dedicated connector jobs and updated-modules shards) on one host JDK:
the lower supported one, JDK 8 today (or 11 after #11545).
2. Keep both JDKs for:
- unit tests;
- engine E2E (`engine-v2-it`, `engine-k8s-it`), where the host JVM runs
engine-side code in-process for some tests;
- transform E2E, which is small.
3. The nightly (`TEST_IN_PR=false`) runs every job on both JDKs. It's a
small change: make the matrix `java` list depend on `inputs.TEST_IN_PR`.
4. Optional: a maintainer label (`full-ci`) restores both JDKs on a PR.
### Expected effect
About 2,500 runner-minutes and about 30 job legs fewer per full-matrix PR
push (roughly 45–50% of IT runner-minutes), and about half the connector legs
on connector-only PRs. This also relieves the fork concurrent-job limit, which
is where most PR wall time goes (jobs waited up to 453 minutes to start).
### Risk and backstop
- A connector bug that only shows with the JDK 11 host client (for example a
driver used in test assertions) would move from PR to nightly. The 66-run
sample found no such case, but the sample is small and from one fork. Before
switching, I can extend it to the #11513 sample (185 runs, 37 forks).
- The nightly must be able to finish for the backstop to be real. Today it
is cancelled by dev pushes (shared concurrency group); #12467 fixes that.
### Questions for dev@
- Which host JDK should the single PR leg use: 8 until #11545 lands, then 11?
- Is the nightly an acceptable backstop for the second JDK on connector E2E?
### Required checks and coordination
- `Build` stays the single required check. This only changes which legs run
inside it on ordinary PRs; the nightly (and a maintainer-applied full-CI
option) keep full coverage, per the boundary set in #11513.
- Complementary to #9900 (merged CI optimization) and #11074 (translation
test dependencies pulling unrelated connector modules into incremental IT
builds); it does not change either.
- This is a proposal for the dev@ discussion requested in #11513; no
implementation before there is agreement.
### Are you willing to submit a PR?
- [x] Yes, once the direction is agreed on dev@.
### Code of Conduct
- [x] I agree to follow this project's [Code of
Conduct](https://www.apache.org/foundation/policies/conduct)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]