aaron-y-chen opened a new pull request, #73998:
URL: https://github.com/apache/airflow/pull/73998

   # Human Summary
   
   Local Airflow development on K8s (`breeze k8s`) is quite heavy on resources 
(CPU/memory). I found that our Kind clusters start two nodes where one is 
enough, and dropping the extra node noticeably reduces resource usage. This 
setup has been really useful for me.
   
   I have used it to develop #69613, #69945 and #72555, and to reproduce a 
scheduler crash while reviewing #70475. Everything has worked well so far.
   
   Please let me know if I've missed anything, for example if there's a 
historical reason for the two-node setup that I'm not aware of. Thanks!
   
   # AI Summary
   
   <details>
   <summary>AI Summary</summary>
   
   ## Why
   
   The K8s test Kind clusters have used a control-plane + worker topology since 
the original Kind migration in 2019 (#5837), and it has never been revisited.
   
   The second node adds no schedulable capacity:
   
   - In a multi-node cluster, kubeadm taints the control-plane `NoSchedule`.
   - None of the pods we deploy tolerate that taint. This covers the chart 
components, executor task pods and KPO test pods.
   - So every workload already runs on the single worker.
   
   The second node still has a cost. `kind load` copies the Airflow image and 
the pinned test images into **every** node, so the control-plane receives a 
~1.8GB copy that nothing reads. Every cluster also has to provision and join 
the extra node.
   
   This PR switches the cluster to a single control-plane node, which is Kind's 
default layout. Kind removes the control-plane taint when a cluster has only 
one node, so the same pods run there instead.
   
   No test coverage is lost. All workload pods already shared one node, so no 
test ever exercised cross-node scheduling.
   
   ## Changes
   
   - `kind-cluster-conf.yaml`: drop the worker node. `extraPortMappings` moves 
to the control-plane.
   - `get_kubernetes_port_numbers()`: the forwarded port was read from a 
hard-coded `nodes[1]`, which raises `IndexError` on the new config. The mapping 
is now looked up by `containerPort`, so the rendered configs of clusters 
created **before** this change keep working too.
   - Unit tests for the port extraction: single-node, legacy two-node and 
missing mapping.
   - Doc sample outputs and workflow comments updated to match.
   
   ## CI impact
   
   Two CI runs make a same-base A/B comparison:
   
   - the CI run of #69459 
([28988414934](https://github.com/apache/airflow/actions/runs/28988414934))
   - the scheduled canary run an hour later 
([28990785885](https://github.com/apache/airflow/actions/runs/28990785885))
   
   Both were built from the same `main` commit (`4d56e6c93d`) on the same 
runner type (`ubuntu-22.04`).
   
   Phase timings from the job logs (Python 3.10, K8s v1.30.13, 
`use-standard-naming=false`). Each cell shows KubernetesExecutor / 
CeleryExecutor:
   
   | phase | two-node (canary) | single-node (#69459) |
   |---|---|---|
   | `kind create cluster` | 42s / 42s | 26s / 24s |
   | `kind load` Airflow image | 81s / 74s | 61s / 60s |
   | pytest | 354s / 277s | 353s / 271s |
   
   The cluster-creation difference is the worker join, which took ~17s. Test 
time is unchanged.
   
   A wider baseline of successful jobs from 13 canary runs (2026-07-05 to 
2026-07-11) gives the same picture:
   
   - In 5 of the 6 `K8S System` jobs, the `Run complete K8S tests` step ran 
below the baseline minimum. The sixth was within range.
   - The Kustomize overlay smoke test times each phase as its own step. Cluster 
creation went from 42–49s to 25s. The image upload went from 71–84s to 61s.
   
   That is roughly 35–50s saved per K8s job. A canary run has 38 K8s jobs, 
about 475 job-minutes in total: 36 `K8S System` jobs, `K8S Lang-SDK` and the 
overlay smoke test. So this saves about 20–30 runner-minutes per canary run.
   
   K8s jobs are not on the PR critical path, so this change does not shorten PR 
feedback time. The larger gain is for local development, shown below.
   
   ## Local benchmark
   
   This was measured on the original branch: a single run on Apple-silicon 
macOS + Docker Desktop, with the same image on both sides, KubernetesExecutor, 
Python 3.10 and K8s v1.30.13.
   
   | metric | two-node | single-node | delta |
   |---|---|---|---|
   | `upload-k8s-image` | 114.1s | 66.4s | **−42%** |
   | containerd disk, all nodes | 15.1 GB | 7.8 GB | **−48%** |
   | memory after deploy, all nodes | ~3.39 GiB | ~2.65 GiB | **−22%** |
   | `kind create cluster` ¹ | 50.5s | 16.8s | **−67%** |
   | `deploy-airflow` | 97.0s | 83.9s | −14% |
   | test results | 57 passed / 3 skipped | 57 passed / 3 skipped | identical |
   
   ¹ Measured with plain `kind create cluster` and the node image cached on 
both sides, to isolate the effect of the topology.
   
   ## Validation
   
   On this branch, rebased onto `main` at `489316fbc7`:
   
   - `uv run --project dev/breeze pytest 
dev/breeze/tests/test_kubernetes_utils.py -xvs`: 3 passed.
   - `prek run --stage pre-commit` on the changed files: passed.
   
   With the same code on #69459:
   
   - CI: all six `K8S System` jobs, `K8S Lang-SDK` and the Kustomize overlay 
smoke test passed.
   - Full manual `breeze k8s` flow (create / configure / upload / deploy / 
tests / delete):
     - KubernetesExecutor: 57 passed / 3 skipped, identical to the two-node 
baseline.
     - CeleryExecutor, which has the most pods (redis + celery workers): 53 
passed / 7 skipped.
   
   ## Notes for reviewers
   
   - The CI run on #69459 covered only the default Python 3.10 / K8s v1.30.13 
combination. The `full tests needed` label would run the full K8s matrix (v1.30 
to v1.35) on single-node clusters before merge.
   - Release branches still read `nodes[1]`. Suppose you create a single-node 
cluster from `main`, then check out a release branch in the same worktree. 
`breeze k8s` on that branch fails with `IndexError`. Running `breeze k8s 
create-cluster --force-recreate-cluster` recovers.
   
   </details>
   
   ---
   
   ##### Was generative AI tooling used to co-author this PR?
   
   - [x] Yes (please specify the tool below)
   
   Generated-by: Claude Code (Opus 5.5) following [the 
guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to