[
https://issues.apache.org/jira/browse/CAMEL-24290?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121260#comment-18121260
]
Claus Ibsen commented on CAMEL-24290:
-------------------------------------
Merged the recovered-flake reporting step to main via
https://github.com/apache/camel/pull/25598 (4.23.0). Leaving this issue open as
the umbrella for further CI work.
_Claude Code on behalf of davsclaus_
> Improve Continuous Integration to support larger workload
> ---------------------------------------------------------
>
> Key: CAMEL-24290
> URL: https://issues.apache.org/jira/browse/CAMEL-24290
> Project: Camel
> Issue Type: Task
> Components: build system
> Affects Versions: 4.21.0
> Reporter: Aurélien Pupier
> Priority: Major
>
> Our 2 CI systems (Jenkins and GitHub actions) are already under heavy loads
> and the planned future will surely increase the load and the need for a more
> stable CI.
> Currently Jenkins is taking between 4 and 5 hours when things are going well.
> Can go up to 20 hours when all streams in //
> On GitHub, it happens regularly that the jobs are queued for several hours.
> We have few full green build per week on Jenkins. (We are at one or two per
> year last year, so already a lot better but not enough to really sustain new
> workloads efficiently).
> Having a CI failing only for real issues will also help all committers to
> react faster
>
> There are several factors for increased, and potential new increase:
> * Increased number PRs with more review, so more cycle
> * Increased number of backports
> * Increased cadence of dependencies release
> * Longer LTS ? , more streams to maintain?
> * Shared GitHub runners with all Apache projects (and several of them are
> facing the same workload increase)
> Several areas identified which worth at least some investigations:
> * stability of Workflows run: it gives more confidence to the build, avoids
> the need to rerun to check that it is "just" flaky or an infra hiccups
> ** develocity help tracking flaky tests (especially flaky tests causing
> failure)
> ** -one last test is failing very often: SjmsConnectionRecoveryTest-
> -https://issues.apache.org/jira/browse/CAMEL-24284- DONE no failure since 1st
> September
> ** several failing less often (every months)
> *** useful Develocity dashboard
> [https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.relativeStartTime=P28D]
> **** beware of tests which failed for good reasons (it caught regressions)
> **** beware of all the tests which were already fixed (there was a huge
> session of fix, so maybe need to wait a month to have a clearer view and/or
> play with the the timeframe)
> **** I filtered on main and excluded JDK 26 for now. Because i have JDK 26
> has tests which are failing for an identified bug and a lot of stability
> improvements were done on main only
> **** based on previous warning, the current curated list is:
> ***** MasterEndpointIT
> [https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.timeZoneId=Europe%2FParis&tests.container=org.apache.camel.component.zookeepermaster.MasterEndpointIT&tests.test=testEndpoint]
> ***** MongoDB:
> [MongoDbBulkWriteOperationsIT|https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.timeZoneId=Europe%2FParis&tests.container=org.apache.camel.component.mongodb.integration.MongoDbBulkWriteOperationIT],
>
> [MongoDbCredentialsFromUriConnectionIT|https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.timeZoneId=Europe%2FParis&tests.container=org.apache.camel.component.mongodb.integration.MongoDbCredentialsFromUriConnectionIT]
> and
> [MongoDbFindOneAndReplaceOperationIT|https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.timeZoneId=Europe%2FParis&tests.container=org.apache.camel.component.mongodb.integration.MongoDbFindOneAndReplaceOperationIT]
> *****
> [SpringFileAntPathMatcherRemoteFileFilterTest|https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.timeZoneId=Europe%2FParis&tests.container=org.apache.camel.itest.ftp.SpringFileAntPathMatcherRemoteFileFilterTest]
> CAMEL-24618
> *****
> [CassandraIdempotentRespotoryIT|https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main%20-maven.environment.jvmVersion:*26*&search.timeZoneId=Europe%2FParis&tests.container=org.apache.camel.processor.idempotent.cassandra.CassandraIdempotentRepositoryIT]
> ***** and certainly a few others currently buried in the list of already
> fixed ones
> ** and also several flaky during the run but passing in less than 3
> attempts, slowing down the build
> ***
> [https://develocity.apache.org/scans/tests?search.query=project:camel%20tag:main&search.relativeStartTime=P28D&search.timeZoneId=Europe%2FParis&tests.sortField=FLAKY]
> ***
> ** -one Jenkins workspace leak to fix due to some jbang tests- [CAMEL-23795]
> DONE
> * modification of test matrixes
> ** testing a single JDK/OS permanently and push to weekly/nightly others on
> jenkins?
> ** -Reduce test matrix on maintenance stream?- DONE
> * Optimize usage of available resources
> ** see if we can build on smaller instances of Jenkins, for Linux it will
> allow to move up from 4 to ~20 instances available.
> ** use several threads (but requires a far better stability and robustness
> of tests and ensure no port conflict)
> https://issues.apache.org/jira/browse/CAMEL-24318
> *** to check also if we can do the pure compile in // with all threads and
> then launch tests as of now (1 threads used although 22 are available on big
> machines)
> **** on GitHub there are already 2 phases but it is using mvnd; [it is
> taking around 18
> minutes|https://github.com/apache/camel/actions/runs/30412528538/job/90451766586#step:9:1]
> **** on jenkins it means to split in different phase
> **** to test for sonar scanner ([currently 1 thread is
> used|https://develocity.apache.org/s/fjldcrmqc7bds] and it is taking between
> around 30 minutes)
> *** on Jenkins, I think we could attempt to provide a specific job for this
> experiment with specific develocity tag so that can filter out/in the results
> easily
> ** also seems 3.5G of Memory is used for the build, maybe we could use more
> when in //
> *** 32G in total on large jenkins machine
> *** but [develocity seems to report only 1.7
> max|https://develocity.apache.org/s/kuigqxgdpicsy/performance/build] is used
> currently
> *** maybe an indication that we could now run on smaller Jenkins node as
> this Memory constraint seems to date from 2018 and it has not changed, good
> chance that jenkins nodes size has been improved since that time
> * several s390x and ppc64le in //, currently [3
> ppc64le|https://jenkins-ccos.apache.org/view/Shared%20-%20ppc64le%20nodes/]
> and [6
> s390x|https://jenkins-ccos.apache.org/view/Shared%20-%20s390x%20nodes/] nodes
> are referenced as shared node but a single one is available at a time
> ** started dicussion here
> [https://the-asf.slack.com/archives/CF6PY5M4N/p1787908008691489]
> ** Bug opened in Jenkins [https://github.com/jenkinsci/jenkins/issues/27318]
> * improve performance of the product/test code itself maybe the trickier one
> but also the most beneficial for our end users
> * use of Develocity build cache (not a big fan of not rebuilding but could
> be something to reuse in some configurations)
> * use of Develocity test selection
> * improve scalpel to reduce even more the scope of PR tests
> * improve [fetch based branch for scalpel change detection on Jenkins which
> is taking almost 5
> minutes|https://github.com/apache/camel/actions/runs/30429114691/job/90501998254#step:3:1]
> * gating PR builds?
> ** have a specific account for AI agents which does not trigger build
> automatically
> ** only a subset triggered automatically then full test suite launched on
> demand for all? for specific accounts?
> * see if budget to provide more github runners and/or jenkins nodes
> * check for performance bottleneck in the build and test process (develocity
> build scan performance can give some ideas)
--
This message was sent by Atlassian Jira
(v8.20.10#820010)