xlsongc commented on code in PR #73926:
URL: https://github.com/apache/airflow/pull/73926#discussion_r4165485707


##########
providers/common/ai/src/airflow/providers/common/ai/operators/document_loader.py:
##########
@@ -452,19 +454,56 @@ def _parse_pdf_stream(self, stream: BinaryIO) -> 
list[dict[str, Any]]:
 
     def _parse_docx_stream(self, stream: BinaryIO) -> list[dict[str, Any]]:
         """
-        Parse a DOCX stream into documents.
+        Parse a DOCX stream into a single document.
 
-        Extracts paragraph text only. Tables, headers, footers, and footnotes
-        are not included. For richer DOCX parsing, plug in a dedicated
-        extraction tool (``Unstructured``, ``docling``) as a custom parser
-        backend.
+        Paragraphs and tables in the document body are extracted in document
+        order. Each table row becomes one "| cell | cell |" line, and a nested
+        table is flattened into its cell. Headers, footers, footnotes, and
+        content controls are not included.
         """
         try:
             from docx import Document
+            from docx.table import Table
         except ImportError as e:
             raise AirflowOptionalProviderFeatureException(e)
 
         doc = Document(stream)
-        paragraphs = [p.text for p in doc.paragraphs if p.text.strip()]
-        text = "\n\n".join(paragraphs)
+        blocks = []
+        for block in doc.iter_inner_content():
+            if isinstance(block, Table):
+                text = "\n".join(f"| {' | '.join(cells)} |" for cells in 
self._get_docx_table_rows(block))
+            else:
+                text = block.text
+            if text.strip():
+                blocks.append(text)
+        text = "\n\n".join(blocks)
         return [{"text": text, "metadata": {}}]
+
+    def _get_docx_table_rows(self, table: Table) -> list[list[str]]:
+        rows = []
+        for row in table.rows:
+            cells: list[str] = []
+            previous_cell = None
+            for cell in row.cells:

Review Comment:
   Done. Rows are padded with `grid_cols_before` / `grid_cols_after` empty 
cells, and `test_docx_grid_before_and_after_keep_columns` builds a real table 
with one `w:gridBefore` row and one `w:gridAfter` row. Also checked with a 
table saved by Word after "Delete cells → Shift cells left", which writes 
`w:gridAfter`.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to