The GitHub Actions job "Required Checks" on texera.git/feat/parquet-source has succeeded. Run started by GitHub user kz930 (triggered by kz930).
Head commit for run: 62cf952843b9c87e983dbb59da26150293e01e7e / kary zheng <[email protected]> feat(operator): read a Parquet file as a source Texera read CSV, JSONL, Arrow and plain text off disk but not Parquet, the format most tables in a data-science workflow are already stored in. Converting one to CSV first loses what the file knew: the column written as an INTEGER came back as text for the schema to guess at again. The file states its own types in a footer, so this source infers nothing. It reads a row group at a time rather than the whole file, which is the thing a columnar format is chosen to avoid, and a column that is a group, a list or a map is refused by name instead of being dropped in silence. A timestamp is read with UTC arithmetic, matching what ArrowUtils means by a Texera TIMESTAMP: the count from the epoch lands on a wall clock, and no zone of the machine's own enters it. Reading it any other way would move the value and part the engine from the script the export writes. No new dependency: parquet-hadoop and parquet-column are already on the classpath under iceberg-parquet, and the reader takes a LocalInputFile so none of Hadoop's own file plumbing is involved. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Report URL: https://github.com/apache/texera/actions/runs/34575308765 With regards, GitHub Actions via GitBox
