xlsongc opened a new pull request, #73926:
URL: https://github.com/apache/airflow/pull/73926
DOCX parsing in `DocumentLoaderOperator` only read `doc.paragraphs`, so
tables were
silently dropped and never reached downstream embeddings. PDF parsing
already keeps
table text.
Body paragraphs and tables are now extracted in document order, one line per
table row:
```text
Quarterly results
| Region | Revenue |
| EMEA | 1.2M |
```
- A cell merged across columns appears once; a cell merged down several rows
repeats
on each row, so every row reads on its own.
- A nested table is flattened into its cell (`/` between cells, `;` between
rows).
- Headers, footers, footnotes and content controls are still not included
(documented).
Notes for reviewers:
- The `docx` extra now requires `python-docx>=1.1.0` for
`iter_inner_content()`.
python-docx is also added to the `dev` group so the tests build real
documents
instead of mocking the module. The two existing mock-based DOCX tests now
use real
documents; their assertions are unchanged.
- Existing DOCX files that contain tables produce different text after this
change,
so re-embedding them yields different vectors.
- Tables are plain XML in DOCX, so the built-in loader handles them without
a heavy
dependency; docling remains the suggestion for richer parsing.
Tested locally: unit tests with python-docx 1.2.0 and 1.1.0, mypy on the
changed module,
prek pre-commit hooks, and the provider docs build. The breeze-based checks
(`mypy-providers`, tests in the CI image) could not run locally, so CI
covers them.
---
##### Was generative AI tooling used to co-author this PR?
- [X] Yes (please specify the tool below)
Generated-by: Claude Code (Opus 5.5) following [the
guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]