L. C. Hsieh created SPARK-59224:
-----------------------------------

             Summary: Add a nightly end-to-end workflow running the gateway on 
kind with a real PySpark client
                 Key: SPARK-59224
                 URL: https://issues.apache.org/jira/browse/SPARK-59224
             Project: Spark
          Issue Type: Sub-task
          Components: Connect
    Affects Versions: connect-gateway-0.1.0
            Reporter: L. C. Hsieh
            Assignee: L. C. Hsieh


The deploy/examples/e2e-smoke walkthrough exercises the whole deployment path —
gateway image, Helm chart, Kubernetes Endpoints discovery, session affinity, 
audit
and metrics — but only by hand. Nothing in CI covered it, so a break in the 
chart,
the Dockerfile or the K8s pool would only surface when someone next ran the
walkthrough manually.

This adds a separate E2E workflow that automates that walkthrough:

  build the gateway image -> create a kind cluster -> load the image -> deploy 
two
  apache/spark:4.0.0 Spark Connect servers -> install the gateway with the Helm
  run the repo's own test/integration/client_smoke.py through a port-forward ->
  assert scg_backend_pool_size and the ExecutePlan counters, and that the audit 
log
  recorded ExecutePlan.

It is deliberately NOT part of the PR-gating CI. Measured locally the 
walkthrough
  takes about 7.5 minutes end to end (image build 5m23s, Spark image pull and 
pod
  readiness 54s, everything else seconds); on a 2-core hosted runner expect 
roughly
  12-20 minutes, dominated by `cargo build --release` and the ~700 MiB Spark 
image
  pull. So it runs nightly at 06:00 UTC and can be triggered by hand from the
  Actions tab whenever the Dockerfile, the chart or the manifests change.
  
  kind and helm are installed with plain curl at pinned versions rather than
  third-party actions, both to stay clear of the ASF GitHub Actions allowlist 
and to
  keep the versions explicit. The only action used is actions/checkout. On 
failure
  the job dumps pod state and gateway/Spark logs; the kind cluster is always 
deleted.

  Verified by running the entire walkthrough locally first: the PySpark client
  returned correct results including a TempView query, which is the meaningful 
check
  that session affinity held (a TempView lives in one driver's memory, so a
  misrouted follow-up RPC would fail it).




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to