sabz90 opened a new pull request, #74392:
URL: https://github.com/apache/airflow/pull/74392
## Summary
Avoid materializing large `DagRun.conf` object graphs in the Execution API
server and task supervisor.
The task-start endpoint now:
1. defers the ORM `DagRun.conf` attribute;
2. selects `CAST(dag_run.conf AS TEXT)` once;
3. returns that text in a dedicated `TIRunContext.dag_run_conf_json` field;
4. leaves `TIRunContext.dag_run.conf` as `None`;
5. carries the string unchanged through the supervisor protocol;
6. decodes it once in the task process immediately before params and template
context construction.
The decoded dictionary is cached on `dag_run.conf`, and the serialized string
is released, so repeated `get_template_context()` calls do not parse it
again.
Related: #74025
## Scope
This intentionally addresses only task-start memory amplification:
- no database type or migration changes;
- no exponent-number persistence changes from #74023/#74093;
- no generic MessagePack large-integer changes from #73712/#73783.
It replaces the memory portion of closed draft #74182 following review
feedback
that the earlier PR mixed persistence correctness and transport optimization.
## Difference from #74048
#74048 first materialized `DagRun.conf` in the Execution API and supervisor,
then serialized it back to JSON before sending it to the task.
This change selects serialized text directly from the database. The API
server
and supervisor never construct the nested dictionary. The task still
materializes it once because existing params, templates, and DAG code require
dictionary behavior.
## Compatibility
- Normal `DagRun` API responses keep the existing dictionary schema.
- Only `TIRunContext` gains the optional `dag_run_conf_json` field.
- Older Execution API versions receive the legacy dictionary through Cadwyn.
- Older foreign SDK supervisor versions receive the legacy dictionary and do
not receive the new field.
- `NULL` configuration remains `None`.
- Empty `{}` configuration is preserved and decoded once.
## Evidence
Synthetic isolated-phase benchmark with an 11,461,903-byte nested payload:
| Phase | Current peak | String peak | Current time | String time |
|---|---:|---:|---:|---:|
| API | 26.74 MB | 17.19 MB | 119.27 ms | 9.77 ms |
| Supervisor | 26.66 MB | 17.19 MB | 229.62 ms | 6.17 ms |
| Task | 40.79 MB | 42.78 MB | 190.00 ms | 133.27 ms |
The task-side cost remains. This draft still needs concurrent end-to-end pod
RSS
and startup-latency evidence before it is ready.
## Validation
Passed in Apache's Linux CI image:
- Execution API `NULL` conf regression;
- Execution API compact-text task-start transport;
- Task SDK supervisor decode;
- one-time task-process materialization and repeated-context reuse;
- older foreign SDK downgrade conversion;
- Ruff on every changed Python file;
- Execution API and supervisor schema-version checks;
- `git diff --check`.
The query uses `selectinload` for consumed asset events so the large JSON
string
is not duplicated once per joined collection row.
---
##### Was generative AI tooling used to co-author this PR?
- [x] Yes — GitHub Copilot (implementation support, test generation, and PR
description drafting). The design and changes were reviewed and validated
by
the contributor.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]