Lee-W commented on code in PR #73925:
URL: https://github.com/apache/airflow/pull/73925#discussion_r4141133591


##########
providers/common/ai/src/airflow/providers/common/ai/toolsets/object_storage.py:
##########
@@ -312,7 +322,13 @@ def _read_file(self, relative: str, *, offset: int | None, 
limit: int | None) ->
                     sample = sample_columnar_file(
                         target, file_format=columnar, 
sample_rows=_SAMPLE_ROWS, max_bytes=self._max_read_bytes
                     )
-                    return _cut(sample, self._max_output_bytes)
+                    # First, so that cutting a long sample never drops it.
+                    shown = min(_SAMPLE_ROWS, sample.total_rows)
+                    header = (
+                        f"Rows: {sample.total_rows}. The schema and the first 
{shown} rows follow; "
+                        f"offset and limit do not apply to 
{columnar.capitalize()} files.\n"
+                    )
+                    return _cut(header + sample.text, self._max_output_bytes)

Review Comment:
   ```suggestion
                           f"offset and limit do not apply to 
{columnar.capitalize()} files."
                       )
                       return _cut(f"{header}\n{sample.text}", 
self._max_output_bytes)
   ```
   
   non-blocking, but i prefer doing something like this. although the 
difference is not noticeable 



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to