found out what the problem was. It turned out that spark was consuming too
much memory and not enough was left for OS. When doing large shuffle writes,
performance is greatly reduced if there is not enough memory left for OS
cache buffer.
We have changed our configuration that spark on workers only uses up to 75%
of available Memory. This can be achieved in spark-env.sh with the
parameter:
SPARK_WORKER_MEMORY=50g (67g available on the system)
we have a small script that checks how much memory there is available and
leaves 75% of it to spark:
MEM_KB=`cat /proc/meminfo | grep MemTotal | awk '{print $2}'`
MEM=$[(MEM_KB * 75 / 100) / 1024]
echo "export SPARK_MEM=${MEM}m" >> /opt/spark/conf/spark-env.sh
echo "export SPARK_WORKER_MEMORY=${MEM}m" >> /opt/spark/conf/spark-env.sh
--
View this message in context:
http://apache-spark-user-list.1001560.n3.nabble.com/Large-shuffle-RDD-tp2649p2692.html
Sent from the Apache Spark User List mailing list archive at Nabble.com.