[
https://issues.apache.org/jira/browse/NIFI-14502?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17947869#comment-17947869
]
MichaelOF commented on NIFI-14502:
----------------------------------
VOILA.... REPRODUCIBLE:
Steps to reproduce:
(when data flow works fine)
1. <NIFI_HOME>/bin/nifi.sh start (No need to start any processors, only Nifi)
2. "POWER OFF" VM, on which Nifi (2.3.0) is installed. Don't "nifi.sh stop"
before.
3. Restart VM.
4. Restart Nifi
Result: Flow files are HANGING FOREVER, as described....
> "Corrupted" Processor Group, Process Group Outbound Policy "Batch Output".
> Flow files are not entering PG
> ---------------------------------------------------------------------------------------------------------
>
> Key: NIFI-14502
> URL: https://issues.apache.org/jira/browse/NIFI-14502
> Project: Apache NiFi
> Issue Type: Bug
> Affects Versions: 2.3.0
> Environment: RHEL 9, standalone install (no docker/K8s). VM on VMware
> Aria Automation
> Reporter: MichaelOF
> Priority: Major
>
> I have a data flow, where 2 processor groups are connected, building a
> sequence.
> PG 1: Single FlowFile Per Node / Batch Output
> PG 2: Single Batch Per Node / Batch Output
> Simple basic Nifi installation. No remote ports etc., no multiple instances.
> Data flow itself is exported as JSON from 1.23.2 (cetic/nifi, helm deploy)
> and imported into a brandnew 2.3.0 standalone Nifi instance. Not sure if
> relevant...
> Data flow is a batch processing, running on demand (only). Rock solid stable
> on 1.23.2. Rock solid on 2.3.0, until this issue happens......
> Issue is the following, [asked in Slack
> already|https://apachenifi.slack.com/archives/C0L9VCD47/p1745337237774879]:
> When issue "appears", Nifi 2.3.0 does NOT recognize anymore, as before, that
> ALL flow files are ready to leave PG 1 / ready to enter PG 2. They keep
> staying queued directly before PG 1's outbound port. Currently "forever", as
> no exp. date set.
> Workaround, found and described in Slack, is to
> # copy "non-receiving-anymore" PG
> # paste it
> # re-wire batch output PG to pasted receiving PG
> # (drop now unnecessary, unconnected original PG)
> I have NO CLUE what might be wrong at that point with PG 2, how it get's
> "corrupted". But as ANY new PG will immediately accept the flow files from
> PG 1, I call it "corrupted". Tried a very long while to narrow down to that
> point.
> What's NEW to me, didn't know this when asking in Slack, is that it seems to
> be related to a VM shutdown and reboot, on VMware Aria side. Which happens
> once a week, during the night from Sunday to Monday. Means from yesterday
> until today. ALTHOUGH the data flow and all of it's processors and PGs are
> OFF most of the time! But used the data flow on Friday, always fine. Tried
> today, issue at once, within first run! Applied "copy/pasted/drop PG 2"
> workaround - again running fine all times...
> NO IDEA what might be the reason....
> I'll to test with a "hard shutdown" of my VM, if I can repoduce this issue,
> and report here.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)