What version of Spark SQL are you running here? I think a lot of your concerns have likely been addressed in more recent versions of the code / documentation. (Spark 1.1 should be published in the next few days)
In particular, for serious applications you should use a HiveContext and HiveQL as this is a much more complete implementation of a SQL Parser. The one in SQL context is only suggested if the Hive dependencies conflict with your application. > 1) spark sql does not support multiple join > This is not true. What problem were you running into? > 2) spark left join: has performance issue > Can you describe your data and query more? > 3) spark sql’s cache table: does not support two-tier query > I'm not sure what you mean here. > 4) spark sql does not support repartition You can repartition SchemaRDDs in the same way as normal RDDs.
