The GitHub Actions job "Check newsfragment PR number" on 
airflow.git/batch-parse-result-persistence has failed.
Run started by GitHub user ephraimbuddy (triggered by ephraimbuddy).

Head commit for run:
19b8afa57db42541dc2adc900c14868f819cf93e / Ephraim Anierobi 
<[email protected]>
Persist a sweep of Dag parse results in one pass

The Dag processor's manager wrote parse results one file at a time. Most of what
that costs is fixed per call rather than per Dag, so a sweep of single-Dag files
paid the fixed cost once per file. Persisting a sweep together pays it once: at 
a
group of eight files, 12.0 statements per file become 4.1, and the call itself
runs about twice as fast.

The saving is that fixed cost spread across the group, so it converges -- around
2.2x however wide the group gets -- and it is worth nothing at the default two
parsing processes, where a sweep only ever holds one file. Deployments running a
wide Dag processor are the ones this is for.

Parse duration is per file, so batching means it can no longer be a single value
for the call; it is carried per Dag instead. Callers persisting one file still
pass one value.

A sweep is split into groups: one bundle each, since the bundle determines the
version and version data the write needs, and each dag_id claimed at most once,
since writing two files that define the same dag_id together would merge them
into one Dag and lose the duplicate warning. Groups are written in the order 
they
are returned, and a file is only held back to a group after the one holding the
file it duplicates, so the same file wins a duplicated dag_id as when each was
written on its own. A group is also capped, because the saving flattens out
quickly while the rows one transaction holds locked do not.

Each group is persisted by its own call, and so in its own transaction:
update_dag_parsing_results_in_db rolls the session back before retrying a 
database
error, which would otherwise discard a group already written alongside it while
its files were still recorded as persisted. A group that fails is retried one 
file
at a time, so one unwritable file does not discard the results of the others, 
and
finished processors are closed whether or not the sweep succeeded.

Persistence moves out of handle_parsing_result, which keeps the stat and metric
handling. That method and persist_parsing_result were released in 3.3 as the 
seams
for deployments forwarding results somewhere other than the metadata DB. Both 
keep
their released contract and are still called once per file, and both are now
deprecated in favour of persist_parsing_results, which is handed a whole group 
at
once. Overriding a deprecated seam costs that deployment its batching and 
nothing
else.

>From @seanmuth's investigation of the Airflow 2-to-3 dag-processor CPU 
>regression:
github.com/seanmuth/af2-af3-scheduler-cpu-repro. This is the per-sweep result
batching that investigation identified as the largest lever.

Report URL: https://github.com/apache/airflow/actions/runs/32146047057

With regards,
GitHub Actions via GitBox


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to