What version of Spark SQL are you running here?  I think a lot of your
concerns have likely been addressed in more recent versions of the code /
documentation.  (Spark 1.1 should be published in the next few days)

In particular, for serious applications you should use a HiveContext and
HiveQL as this is a much more complete implementation of a SQL Parser.  The
one in SQL context is only suggested if the Hive dependencies conflict with
your application.


> 1)  spark sql does not support multiple join
>

This is not true.  What problem were you running into?


> 2)  spark left join: has performance issue
>

Can you describe your data and query more?


> 3)  spark sql’s cache table: does not support two-tier query
>

I'm not sure what you mean here.


> 4)  spark sql does not support repartition


You can repartition SchemaRDDs in the same way as normal RDDs.

Reply via email to