[ 
https://issues.apache.org/jira/browse/NIFI-16339?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Alexander Bij updated NIFI-16339:
---------------------------------
    Description: 
I'm seeing build failures caused by timing-sensitive race conditions in tests. 
They
intermittently fail CI on unrelated pull requests, forcing maintainers to 
re-run jobs.

h3. Affected tests

*PythonNarDeletionDuringInitIT.testNarReuploadAfterForceDeleteDuringInit*
After a NAR is re-uploaded and reaches INSTALLED, the processor type may not be 
immediately
visible. The test can fail if it asserts the type is available too early.

*TestStandardProcessScheduler.validateNeverEnablingServiceCanStillBeDisabled*
A controller service may transition from DISABLING to DISABLED faster than the 
test expects.
The test currently assumes the service will still be DISABLING at assertion 
time.

h3. Impact

These are flaky timing issues rather than functional regressions. The tests 
should be made
more tolerant of valid asynchronous state transitions so the build is stable.

h3. Expected behavior

- Processor type discovery should be allowed a short time to catch up after NAR 
re-upload.
- The controller service test should accept valid state transitions that happen 
quickly
  (either DISABLING or DISABLED).

h3. Out of scope — separate flaky tests

Multiple clustered/system tests intermittently hit their 5-minute @Timeout 
during node
startup on CI (worst on ubuntu-24.04 Java 25). The failing test rotates 
run-to-run, e.g.:

- AutoResumeStateClusteredIT.testRestartWithAutoResumeStateFalse
- 
OffloadContentClaimTruncationIT.testOffloadedFlowFileContentNotPrematurelyTruncated
- FlowSynchronizationIT.testReconnectAddsProcessor
- ClusteredConnectorTroubleshootingIT.*
- ControllerServiceStateIT.testLocalClusterState

Because there is no single root cause in test logic (only shared timeout 
pressure), these
are tracked separately as CI/runner performance flakiness, not in this ticket.

  was:
I’m seeing build failures caused by timing-sensitive race conditions in tests.

*Affected tests*
 - {{PythonNarDeletionDuringInitIT.testNarReuploadAfterForceDeleteDuringInit}}
 -- After a NAR is re-uploaded and reaches INSTALLED, the processor type may 
not be immediately visible.
 -- The test can fail if it asserts the type is available too early.

 - 
{{TestStandardProcessScheduler.validateNeverEnablingServiceCanStillBeDisabled}}
 -- A controller service may transition from DISABLING to DISABLED faster than 
the test expects.
 -- The test currently assumes the service will still be DISABLING at assertion 
time.

*Impact*
These appear to be flaky timing issues rather than functional regressions. The 
tests should be made more tolerant of valid asynchronous state transitions so 
the build is stable.

*Expected behavior*
Processor type discovery should be allowed a short time to catch up after NAR 
re-upload.
The controller service test should accept valid terminal state transitions that 
happen quickly.


> Stabilize Python NAR reupload and controller service scheduling race 
> conditions in tests
> ----------------------------------------------------------------------------------------
>
>                 Key: NIFI-16339
>                 URL: https://issues.apache.org/jira/browse/NIFI-16339
>             Project: Apache NiFi
>          Issue Type: Test
>          Components: Core Framework
>    Affects Versions: 2.11.0
>            Reporter: Alexander Bij
>            Priority: Minor
>              Labels: test-flaky
>
> I'm seeing build failures caused by timing-sensitive race conditions in 
> tests. They
> intermittently fail CI on unrelated pull requests, forcing maintainers to 
> re-run jobs.
> h3. Affected tests
> *PythonNarDeletionDuringInitIT.testNarReuploadAfterForceDeleteDuringInit*
> After a NAR is re-uploaded and reaches INSTALLED, the processor type may not 
> be immediately
> visible. The test can fail if it asserts the type is available too early.
> *TestStandardProcessScheduler.validateNeverEnablingServiceCanStillBeDisabled*
> A controller service may transition from DISABLING to DISABLED faster than 
> the test expects.
> The test currently assumes the service will still be DISABLING at assertion 
> time.
> h3. Impact
> These are flaky timing issues rather than functional regressions. The tests 
> should be made
> more tolerant of valid asynchronous state transitions so the build is stable.
> h3. Expected behavior
> - Processor type discovery should be allowed a short time to catch up after 
> NAR re-upload.
> - The controller service test should accept valid state transitions that 
> happen quickly
>   (either DISABLING or DISABLED).
> h3. Out of scope — separate flaky tests
> Multiple clustered/system tests intermittently hit their 5-minute @Timeout 
> during node
> startup on CI (worst on ubuntu-24.04 Java 25). The failing test rotates 
> run-to-run, e.g.:
> - AutoResumeStateClusteredIT.testRestartWithAutoResumeStateFalse
> - 
> OffloadContentClaimTruncationIT.testOffloadedFlowFileContentNotPrematurelyTruncated
> - FlowSynchronizationIT.testReconnectAddsProcessor
> - ClusteredConnectorTroubleshootingIT.*
> - ControllerServiceStateIT.testLocalClusterState
> Because there is no single root cause in test logic (only shared timeout 
> pressure), these
> are tracked separately as CI/runner performance flakiness, not in this ticket.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to