aaron-y-chen opened a new pull request, #73998: URL: https://github.com/apache/airflow/pull/73998
# Human Summary Local Airflow development on K8s (`breeze k8s`) is quite heavy on resources (CPU/memory). I found that our Kind clusters start two nodes where one is enough, and dropping the extra node noticeably reduces resource usage. This setup has been really useful for me. I have used it to develop #69613, #69945 and #72555, and to reproduce a scheduler crash while reviewing #70475. Everything has worked well so far. Please let me know if I've missed anything, for example if there's a historical reason for the two-node setup that I'm not aware of. Thanks! # AI Summary <details> <summary>AI Summary</summary> ## Why The K8s test Kind clusters have used a control-plane + worker topology since the original Kind migration in 2019 (#5837), and it has never been revisited. The second node adds no schedulable capacity: - In a multi-node cluster, kubeadm taints the control-plane `NoSchedule`. - None of the pods we deploy tolerate that taint. This covers the chart components, executor task pods and KPO test pods. - So every workload already runs on the single worker. The second node still has a cost. `kind load` copies the Airflow image and the pinned test images into **every** node, so the control-plane receives a ~1.8GB copy that nothing reads. Every cluster also has to provision and join the extra node. This PR switches the cluster to a single control-plane node, which is Kind's default layout. Kind removes the control-plane taint when a cluster has only one node, so the same pods run there instead. No test coverage is lost. All workload pods already shared one node, so no test ever exercised cross-node scheduling. ## Changes - `kind-cluster-conf.yaml`: drop the worker node. `extraPortMappings` moves to the control-plane. - `get_kubernetes_port_numbers()`: the forwarded port was read from a hard-coded `nodes[1]`, which raises `IndexError` on the new config. The mapping is now looked up by `containerPort`, so the rendered configs of clusters created **before** this change keep working too. - Unit tests for the port extraction: single-node, legacy two-node and missing mapping. - Doc sample outputs and workflow comments updated to match. ## CI impact Two CI runs make a same-base A/B comparison: - the CI run of #69459 ([28988414934](https://github.com/apache/airflow/actions/runs/28988414934)) - the scheduled canary run an hour later ([28990785885](https://github.com/apache/airflow/actions/runs/28990785885)) Both were built from the same `main` commit (`4d56e6c93d`) on the same runner type (`ubuntu-22.04`). Phase timings from the job logs (Python 3.10, K8s v1.30.13, `use-standard-naming=false`). Each cell shows KubernetesExecutor / CeleryExecutor: | phase | two-node (canary) | single-node (#69459) | |---|---|---| | `kind create cluster` | 42s / 42s | 26s / 24s | | `kind load` Airflow image | 81s / 74s | 61s / 60s | | pytest | 354s / 277s | 353s / 271s | The cluster-creation difference is the worker join, which took ~17s. Test time is unchanged. A wider baseline of successful jobs from 13 canary runs (2026-07-05 to 2026-07-11) gives the same picture: - In 5 of the 6 `K8S System` jobs, the `Run complete K8S tests` step ran below the baseline minimum. The sixth was within range. - The Kustomize overlay smoke test times each phase as its own step. Cluster creation went from 42–49s to 25s. The image upload went from 71–84s to 61s. That is roughly 35–50s saved per K8s job. A canary run has 38 K8s jobs, about 475 job-minutes in total: 36 `K8S System` jobs, `K8S Lang-SDK` and the overlay smoke test. So this saves about 20–30 runner-minutes per canary run. K8s jobs are not on the PR critical path, so this change does not shorten PR feedback time. The larger gain is for local development, shown below. ## Local benchmark This was measured on the original branch: a single run on Apple-silicon macOS + Docker Desktop, with the same image on both sides, KubernetesExecutor, Python 3.10 and K8s v1.30.13. | metric | two-node | single-node | delta | |---|---|---|---| | `upload-k8s-image` | 114.1s | 66.4s | **−42%** | | containerd disk, all nodes | 15.1 GB | 7.8 GB | **−48%** | | memory after deploy, all nodes | ~3.39 GiB | ~2.65 GiB | **−22%** | | `kind create cluster` ¹ | 50.5s | 16.8s | **−67%** | | `deploy-airflow` | 97.0s | 83.9s | −14% | | test results | 57 passed / 3 skipped | 57 passed / 3 skipped | identical | ¹ Measured with plain `kind create cluster` and the node image cached on both sides, to isolate the effect of the topology. ## Validation On this branch, rebased onto `main` at `489316fbc7`: - `uv run --project dev/breeze pytest dev/breeze/tests/test_kubernetes_utils.py -xvs`: 3 passed. - `prek run --stage pre-commit` on the changed files: passed. With the same code on #69459: - CI: all six `K8S System` jobs, `K8S Lang-SDK` and the Kustomize overlay smoke test passed. - Full manual `breeze k8s` flow (create / configure / upload / deploy / tests / delete): - KubernetesExecutor: 57 passed / 3 skipped, identical to the two-node baseline. - CeleryExecutor, which has the most pods (redis + celery workers): 53 passed / 7 skipped. ## Notes for reviewers - The CI run on #69459 covered only the default Python 3.10 / K8s v1.30.13 combination. The `full tests needed` label would run the full K8s matrix (v1.30 to v1.35) on single-node clusters before merge. - Release branches still read `nodes[1]`. Suppose you create a single-node cluster from `main`, then check out a release branch in the same worktree. `breeze k8s` on that branch fails with `IndexError`. Running `breeze k8s create-cluster --force-recreate-cluster` recovers. </details> --- ##### Was generative AI tooling used to co-author this PR? - [x] Yes (please specify the tool below) Generated-by: Claude Code (Opus 5.5) following [the guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
