Hey Mohit, Behind the pyspark.SparkContext, there is SparkContext in JVM, so the overhead of creating a SparkContext is pretty high. Also, during starting and stopping an SparkContext, there are lots of things need to setup and release, maybe there some corner cases make it not so solid enough.
So, It's better to re-use only one SparkContex for all the jobs in one Python scripts. In the meantime, you will have an WebUI to see all the histories. Davies On Fri, Jul 25, 2014 at 10:21 AM, Mohit Jaggi <[email protected]> wrote: > Folks, > I had some pyspark code which used to hang with no useful debug logs. It got > fixed when I changed my code to keep the sparkcontext forever instead of > stopping it and then creating another one later. Is this a bug or expected > behavior? > > Mohit.
