[ 
https://issues.apache.org/jira/browse/FLINK-4193?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15371524#comment-15371524
 ] 

Gyula Fora commented on FLINK-4193:
-----------------------------------

It happened multiple times in a row on the same cluster but on different 
machines. But it does not occur always. The total parallelism that we were 
deploying was about 200. 

Basically what we did was to run a script that deploys 5 jobs one ofter the 
other with 10 sec in between. It usually crashed somewhere in the middle.

> Task manager JVM crashes while deploying cancelling jobs
> --------------------------------------------------------
>
>                 Key: FLINK-4193
>                 URL: https://issues.apache.org/jira/browse/FLINK-4193
>             Project: Flink
>          Issue Type: Bug
>          Components: Streaming, TaskManager
>            Reporter: Gyula Fora
>            Priority: Critical
>
> We have observed several TM crashes while deploying larger stateful streaming 
> jobs that use the RocksDB state backend.
> As the JVMs crash the logs don't show anything but I have uploaded all the 
> info I have got from the standard output.
> This indicates some GC and possibly some RocksDB issues underneath but we 
> could not really figure out much more.
> GC segfault
> https://gist.github.com/gyfora/9e56d4a0d4fc285a8d838e1b281ae125
> Other crashes (maybe rocks related)
> https://gist.github.com/gyfora/525c67c747873f0ff2ff2ed1682efefa
> https://gist.github.com/gyfora/b93611fde87b1f2516eeaf6bfbe8d818
> The third link shows 2 issues that happened in parallel...



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Reply via email to