I think you need to look into how to run mpi runs with the best binding strategy. --bind-to none can in certain cases really fuck up your core distribution (when you have as many sockets as you have). You really want to do 4 core binding per socket to maximize locality of your cores.
Please refer to your MPI mailing list, this is not a question directly related to SIESTA. Let me note that libraries you compile siesta with, have a huge inflict on performance, so asking "why is it slow?" and providing a fdf file might not give any clue (unless you have specific questions to the parallel settings of siesta), but in this case you have a very low BlockSize, try and let siesta find its own, or play a little with it (if you have optimized it then fine ;) ). 2014-12-14 11:01 GMT+00:00 Seyed Mohammad Tabatabaei <[email protected]>: > > Dear all, > > I have a 64-core system. These 64 cores are actually four 16-core > cpus. I also have 128 GB of RAM. I have many SIESTA runs so I run each > of them on 4 cores with the following command: > > $ mpirun --bind-to none -np 4 siesta.3.2 < MoS2.fdf > MoS2.out > > The parallel compilation is correct. Currently I have run the > following .fdf file and it has taken a long time on 4 cores since ">> > Start of run: 6-DEC-2014 16:38:02" and it has not yet finished. Any > helps for speeding up my calculations through changing the .fdf file > is highly appreciated. I have attached my .fdf file. > > Bests, > Mohammad > -- Kind regards Nick
