[
https://issues.apache.org/jira/browse/SPARK-59224?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
L. C. Hsieh updated SPARK-59224:
--------------------------------
Summary: Add an end-to-end workflow running the gateway on kind with a real
PySpark client (was: Add a nightly end-to-end workflow running the gateway on
kind with a real PySpark client)
> Add an end-to-end workflow running the gateway on kind with a real PySpark
> client
> ---------------------------------------------------------------------------------
>
> Key: SPARK-59224
> URL: https://issues.apache.org/jira/browse/SPARK-59224
> Project: Spark
> Issue Type: Sub-task
> Components: Connect
> Affects Versions: connect-gateway-0.1.0
> Reporter: L. C. Hsieh
> Assignee: L. C. Hsieh
> Priority: Major
>
> The deploy/examples/e2e-smoke walkthrough exercises the whole deployment path
> —
> gateway image, Helm chart, Kubernetes Endpoints discovery, session affinity,
> audit
> and metrics — but only by hand. Nothing in CI covered it, so a break in the
> chart,
> the Dockerfile or the K8s pool would only surface when someone next ran the
> walkthrough manually.
> This adds a separate E2E workflow that automates that walkthrough:
> build the gateway image -> create a kind cluster -> load the image ->
> deploy two
> apache/spark:4.0.0 Spark Connect servers -> install the gateway with the
> Helm
> run the repo's own test/integration/client_smoke.py through a port-forward
> ->
> assert scg_backend_pool_size and the ExecutePlan counters, and that the
> audit log
> recorded ExecutePlan.
> It is deliberately NOT part of the PR-gating CI. Measured locally the
> walkthrough
> takes about 7.5 minutes end to end (image build 5m23s, Spark image pull and
> pod
> readiness 54s, everything else seconds); on a 2-core hosted runner expect
> roughly
> 12-20 minutes, dominated by `cargo build --release` and the ~700 MiB Spark
> image
> pull. So it runs nightly at 06:00 UTC and can be triggered by hand from the
> Actions tab whenever the Dockerfile, the chart or the manifests change.
>
> kind and helm are installed with plain curl at pinned versions rather than
> third-party actions, both to stay clear of the ASF GitHub Actions allowlist
> and to
> keep the versions explicit. The only action used is actions/checkout. On
> failure
> the job dumps pod state and gateway/Spark logs; the kind cluster is always
> deleted.
> Verified by running the entire walkthrough locally first: the PySpark client
> returned correct results including a TempView query, which is the
> meaningful check
> that session affinity held (a TempView lives in one driver's memory, so a
> misrouted follow-up RPC would fail it).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]