Hi Cheng Pan,

Yes, I'd seen SPARK-54002 / #52706; nice addition, and I see it as
complementary rather than overlapping. It gives an SQL CLI over Connect via
BeeLine (the Hive sqlline client) + a JDBC driver; the proposal here is
native `spark-sql --remote sc://...` the actual spark-sql command on the
Spark Connect client directly (as spark-shell/pyspark --remote already do),
no JDBC layer.

Why native still seems worth it alongside BeeLine:
- UX continuity: spark-sql users/scripts just add --remote, no tool switch
for an edge-node -> remote-compute migration.
- Consistency: spark-shell/spark-submit/pyspark all have native --remote
under SPARK-49194; spark-sql is the only interactive entry point that
doesn't.
- Minimal, Hive-free stack: no JDBC driver or Hive-derived client.
Aligned with the enterprises, such as ours, is moving off of Hive-coupled
tooling that Connect is about.

Does that make sense? Also, I've moved the PR to ready-for-review. Keen on
your view.

- PR (ready for review): https://github.com/apache/spark/pull/57383
- JIRA: SPARK-58227 <https://issues.apache.org/jira/browse/SPARK-58227>

Thanks,
Dhruv Pratap (Netflix)
Apache JIRA: dhruvpratap


On Mon, Jul 20, 2026 at 8:25 PM Cheng Pan <[email protected]> wrote:

> Just wondering if you have noticed SPARK-54002 (
> https://github.com/apache/spark/pull/52706), which makes BeeLine support
> the Spark Connect JDBC driver and provides an SQL CLI to connect to the
> Connect Server, providing a similar experience to spark-sql.
>
> Thanks,
> Cheng Pan
>
>
>
> On Jul 21, 2026, at 00:50, Dhruv Pratap via dev <[email protected]>
> wrote:
>
> Hi all,
>
> The Spark Connect entry-point umbrella (SPARK-49194) has landed Connect
> support for spark-shell (SPARK-48936), spark-submit (SPARK-48960), and
> pyspark. The one interactive entry point still missing it is the spark-sql
> CLI: there is currently no way to run spark-sql against a Connect server
> (spark-sql --remote sc://...).
>
> Before writing any code, two questions for the Connect maintainers:
>
> 1. Tracking: should this be a sub-task under SPARK-49194, or a new ticket?
> The umbrella lists pyspark/spark-shell/spark-submit and its sub-tasks
> are all resolved, so I want to confirm the right home.
> 2. Process: does this warrant a SPIP, or is it an incremental improvement
> completing the already-accepted Connect pattern? It is user-facing but adds
> no new public API (details below).
>
> Proposed approach (kept minimal):
> - A new SparkConnectSQLCLIDriver in sql/connect/client/jvm, selected
> on --remote / spark.api.mode=connect, mirroring how ConnectRepl
> backs spark-shell --remote (SPARK-48936).
> - Builds a remote session via SparkSession.builder().client(...),
> runs statements via spark.sql(...), and formats results client-side.
> - No change to the existing SparkSQLCLIDriver / sql/hive-thriftserver
> path; the legacy driver stays the default. No existing public API changes.
> - MVP first (-e/-f + interactive loop + basic output); follow-ups for
> full Hive-style output parity.
>
> I have a short design write-up I can share/attach to the JIRA.
>
> This is complementary to the draft "Add Unified Spark CLI" PR (#57313, cc
> @nchammas): that unifies the bin/ scripts and its sql subcommand currently
> wraps the existing spark-sql shell, so a Connect-native spark-sql would
> slot under it cleanly. Happy to coordinate to avoid duplicate CLI churn.
>
> Thanks,
> Dhruv Pratap (Netflix)
> Apache JIRA: dhruvpratap
> [email protected]  ·  [email protected]
>
>
>

Reply via email to