potiuk commented on PR #74190:
URL: https://github.com/apache/airflow/pull/74190#issuecomment-5985532591

   On the open question of why this got worse with the Python 3.11 image: I 
traced it, and it comes from `google-api-core` hitting a change in 
`importlib.metadata` between 3.10 and 3.11. It isn't the interpreter or missing 
bytecode.
   
   Measured in the published `apache/airflow:3.1.8-python3.10` and 
`-python3.11` images, which the compat job builds on:
   - Same work on both sides: neither image ships `.pyc` files, both 
interpreters are built the same way (PGO + LTO, `-O3`) and run equally fast, 
and collector creation loads the same ~1,070 modules with the same Google, 
protobuf (`upb`) and grpc versions.
   - Creating the hook lineage collector on a warm run takes 0.74s on 3.11 vs 
0.41s on 3.10, and 2.28s vs 1.48s cold. Almost all of the extra time is in the 
`__init__` of `google.api_core` (51 → 208 ms) and of 
`google.cloud.secretmanager_v1` (51 → 204 ms).
   - Both run `check_python_version()` when imported. In `google-api-core` 
2.30.0, the version in the 3.1.8 image, that function always calls 
`importlib.metadata.packages_distributions()`, uncached, just to build a 
package name for a possible warning. The two calls made during collector 
creation take 0.41s on 3.11 vs 0.10s on 3.10. That covers nearly all of the gap.
   
   Why `packages_distributions()` is slower on 3.11:
   - On 3.10 it reads only `top_level.txt` for each distribution and skips any 
that don't have one.
   - On 3.11, if `top_level.txt` is missing, it falls back to 
`_top_level_inferred()`. That loads `dist.files`, which means parsing the whole 
`RECORD` file.
   - `top_level.txt` is written only by setuptools-built wheels. In the 3.11 
image, 160 of 431 distributions don't have one (flit/hatch/maturin wheels, 
`apache-airflow-core` itself, litellm, pandas, scipy, openai and others), so 
3.11 parses about 27,000 extra `RECORD` rows on every call. It finds 334 import 
names instead of 228.
   - This came from importlib_metadata v4.7.0 (#330: *"In 
`packages_distributions`, now infer top-level names from `.files()` when a 
`top-level.txt` (Setuptools-specific metadata) is not present"*). It reached 
CPython in 3.11.0a4, via bpo-44893 / python/cpython#30150, which *"Syncs with 
importlib_metadata 4.8.1"*. Python 3.10's stdlib is still at importlib_metadata 
4.6.
   
   A sub-second cost locally becomes the ~23s per event seen in CI once it runs 
in every forked event child on runners capped at about 1.2 CPU.
   
   A newer `google-api-core` helps only partly. `main` pins 2.38.0, which 
caches `packages_distributions()`, so calls drop from 2 to 1. But each forked 
child is a fresh process: in my test the remaining call still took 0.20s vs 
0.05s on 3.10, and the compat image stays on 2.30.0 anyway.
   
   Beyond fixing CI, this is a nice optimisation in its own right. Most tasks 
never report hook lineage, yet on Airflow 2.11–3.1 every OpenLineage event 
(START and COMPLETE, for every task) paid for building an empty collector. That 
meant importing the asset URI handlers of every installed provider and the 
heavy client libraries they pull in, only to find nothing and discard it. With 
this change that work happens only when a hook actually collected lineage. 
Every deployment running OpenLineage on those versions gets cheaper, faster 
event emission, on any Python version: per your capped measurement, the 
fork-to-emit gap dropped from about 22.5s to 0.11s. It also makes OpenLineage 
less sensitive to how many providers are installed and to changes in 
third-party import-time behaviour, like the `google-api-core` / 
`importlib.metadata` combination above.
   
   ---
   Drafted-by: Claude Code (Opus 5.5); reviewed by @potiuk before posting
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to