AlKor13 commented on issue #73525:
URL: https://github.com/apache/airflow/issues/73525#issuecomment-5847425867

   Still reproducible on `main` @ `f863982`, and the collection path has the 
same defect in two more places than the one you named — so fixing the script 
alone would move the failure rather than remove it.
   
   `scripts/ci/prek/update_providers_dependencies.py`, five `Path.read_text()` 
/ `write_text()` calls with no `encoding=`:
   
   ```
   L73   tomllib.loads(pyproject_toml_file_path.read_text())
   L98   yaml.safe_load(provider_yaml_file.read_text())
   L230  DEPENDENCIES_JSON_FILE_PATH.read_text() ...
   L233  old_content = DEPENDENCIES_JSON_FILE_PATH.read_text() ...
   L235  DEPENDENCIES_JSON_FILE_PATH.write_text(new_content)
   ```
   
   `devel-common/src/tests_common/pytest_plugin.py`, which is what invokes it 
during collection, three more:
   
   ```
   L169  AIRFLOW_PYPROJECT_TOML_FILE_PATH.read_text().splitlines()
   L201  PROVIDER_DEPENDENCIES_JSON_HASH_PATH.read_text()
   L203  
PROVIDER_DEPENDENCIES_JSON_HASH_PATH.write_text(calculated_provider_deps_hash)
   ```
   
   L169 reads the root `pyproject.toml` in the locale code page before the 
script is ever called, so on cp1252 a single non-ASCII byte anywhere in that 
file fails collection with the same `UnicodeDecodeError` even after the script 
is fixed. L235/L203 are the write side: on a non-UTF-8 machine they write 
`generated/provider_dependencies.json` and its hash file in the ANSI code page, 
which is then committed or compared against a UTF-8 original.
   
   A static pass over `scripts/`, `dev/`, `devel-common/` and 
`airflow-core/src/` (sparse checkout, so not the whole repo) finds **1,102** 
places of this shape: 844 `read_text`/`write_text`, 166 `open()` in text mode, 
77 `subprocess(text=True)`, 15 `shutil.rmtree()` without `onexc=`, concentrated 
in `scripts/` (557) and `dev/` (455).
   
   Two notes on scoping the fix:
   
   **A repo-wide `encoding="utf-8"` is not busywork that 3.15 will undo.** [PEP 
686](https://peps.python.org/pep-0686/) makes UTF-8 the default in Python 3.15, 
but Airflow supports 3.10+, so every user below 3.15 keeps the bug; naming the 
encoding is what the 3.15 migration would leave behind anyway.
   
   **The 77 `subprocess(text=True)` call sites are a different fix from the 
file ones.** Windows has two default code pages at once — on this machine 
cp1251 (what `text=True` decodes with) and cp866 (what a console child writes) 
— so UTF-8 is correct for a Python child and wrong for `git` or `uv`; after 
3.15 those turn from silent mojibake into an exception.
   
   Disclosure: the inventory came from `winseam audit`, a tool I wrote after 
hitting this repeatedly ([repo](https://github.com/AlKor13/winseam)); it is a 
static pass, so it runs on Linux CI. Happy to send a PR for the eight lines in 
the collection path if that is useful — that is the part that unblocks Windows 
test runs.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to