[
https://issues.apache.org/jira/browse/NIFI-16339?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Alexander Bij updated NIFI-16339:
---------------------------------
Description:
I'm seeing build failures caused by timing-sensitive race conditions in tests.
They
intermittently fail CI on unrelated pull requests, forcing maintainers to
re-run jobs.
h3. Affected tests
*PythonNarDeletionDuringInitIT.testNarReuploadAfterForceDeleteDuringInit*
After a NAR is re-uploaded and reaches INSTALLED, the processor type may not be
immediately
visible. The test can fail if it asserts the type is available too early.
*TestStandardProcessScheduler.validateNeverEnablingServiceCanStillBeDisabled*
A controller service may transition from DISABLING to DISABLED faster than the
test expects.
The test currently assumes the service will still be DISABLING at assertion
time.
h3. Impact
These are flaky timing issues rather than functional regressions. The tests
should be made
more tolerant of valid asynchronous state transitions so the build is stable.
h3. Expected behavior
- Processor type discovery should be allowed a short time to catch up after NAR
re-upload.
- The controller service test should accept valid state transitions that happen
quickly
(either DISABLING or DISABLED).
h3. Out of scope ā separate flaky tests
Multiple clustered/system tests intermittently hit their 5-minute @Timeout
during node
startup on CI (worst on ubuntu-24.04 Java 25). The failing test rotates
run-to-run, e.g.:
- AutoResumeStateClusteredIT.testRestartWithAutoResumeStateFalse
-
OffloadContentClaimTruncationIT.testOffloadedFlowFileContentNotPrematurelyTruncated
- FlowSynchronizationIT.testReconnectAddsProcessor
- ClusteredConnectorTroubleshootingIT.*
- ControllerServiceStateIT.testLocalClusterState
Because there is no single root cause in test logic (only shared timeout
pressure), these
are tracked separately as CI/runner performance flakiness, not in this ticket.
was:
Iām seeing build failures caused by timing-sensitive race conditions in tests.
*Affected tests*
- {{PythonNarDeletionDuringInitIT.testNarReuploadAfterForceDeleteDuringInit}}
-- After a NAR is re-uploaded and reaches INSTALLED, the processor type may
not be immediately visible.
-- The test can fail if it asserts the type is available too early.
-
{{TestStandardProcessScheduler.validateNeverEnablingServiceCanStillBeDisabled}}
-- A controller service may transition from DISABLING to DISABLED faster than
the test expects.
-- The test currently assumes the service will still be DISABLING at assertion
time.
*Impact*
These appear to be flaky timing issues rather than functional regressions. The
tests should be made more tolerant of valid asynchronous state transitions so
the build is stable.
*Expected behavior*
Processor type discovery should be allowed a short time to catch up after NAR
re-upload.
The controller service test should accept valid terminal state transitions that
happen quickly.
> Stabilize Python NAR reupload and controller service scheduling race
> conditions in tests
> ----------------------------------------------------------------------------------------
>
> Key: NIFI-16339
> URL: https://issues.apache.org/jira/browse/NIFI-16339
> Project: Apache NiFi
> Issue Type: Test
> Components: Core Framework
> Affects Versions: 2.11.0
> Reporter: Alexander Bij
> Priority: Minor
> Labels: test-flaky
>
> I'm seeing build failures caused by timing-sensitive race conditions in
> tests. They
> intermittently fail CI on unrelated pull requests, forcing maintainers to
> re-run jobs.
> h3. Affected tests
> *PythonNarDeletionDuringInitIT.testNarReuploadAfterForceDeleteDuringInit*
> After a NAR is re-uploaded and reaches INSTALLED, the processor type may not
> be immediately
> visible. The test can fail if it asserts the type is available too early.
> *TestStandardProcessScheduler.validateNeverEnablingServiceCanStillBeDisabled*
> A controller service may transition from DISABLING to DISABLED faster than
> the test expects.
> The test currently assumes the service will still be DISABLING at assertion
> time.
> h3. Impact
> These are flaky timing issues rather than functional regressions. The tests
> should be made
> more tolerant of valid asynchronous state transitions so the build is stable.
> h3. Expected behavior
> - Processor type discovery should be allowed a short time to catch up after
> NAR re-upload.
> - The controller service test should accept valid state transitions that
> happen quickly
> (either DISABLING or DISABLED).
> h3. Out of scope ā separate flaky tests
> Multiple clustered/system tests intermittently hit their 5-minute @Timeout
> during node
> startup on CI (worst on ubuntu-24.04 Java 25). The failing test rotates
> run-to-run, e.g.:
> - AutoResumeStateClusteredIT.testRestartWithAutoResumeStateFalse
> -
> OffloadContentClaimTruncationIT.testOffloadedFlowFileContentNotPrematurelyTruncated
> - FlowSynchronizationIT.testReconnectAddsProcessor
> - ClusteredConnectorTroubleshootingIT.*
> - ControllerServiceStateIT.testLocalClusterState
> Because there is no single root cause in test logic (only shared timeout
> pressure), these
> are tracked separately as CI/runner performance flakiness, not in this ticket.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)