The GitHub Actions job "Required Checks" on texera.git/feat/parquet-source has 
succeeded.
Run started by GitHub user kz930 (triggered by kz930).

Head commit for run:
62cf952843b9c87e983dbb59da26150293e01e7e / kary zheng <[email protected]>
feat(operator): read a Parquet file as a source

Texera read CSV, JSONL, Arrow and plain text off disk but not Parquet,
the format most tables in a data-science workflow are already stored in.
Converting one to CSV first loses what the file knew: the column written
as an INTEGER came back as text for the schema to guess at again.

The file states its own types in a footer, so this source infers nothing.
It reads a row group at a time rather than the whole file, which is the
thing a columnar format is chosen to avoid, and a column that is a group,
a list or a map is refused by name instead of being dropped in silence.

A timestamp is read with UTC arithmetic, matching what ArrowUtils means by
a Texera TIMESTAMP: the count from the epoch lands on a wall clock, and no
zone of the machine's own enters it. Reading it any other way would move
the value and part the engine from the script the export writes.

No new dependency: parquet-hadoop and parquet-column are already on the
classpath under iceberg-parquet, and the reader takes a LocalInputFile so
none of Hadoop's own file plumbing is involved.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

Report URL: https://github.com/apache/texera/actions/runs/34575308765

With regards,
GitHub Actions via GitBox

Reply via email to