The GitHub Actions job "Backport Approval Check" on 
texera.git/feat/standalone-sources has succeeded.
Run started by GitHub user kz930 (triggered by kz930).

Head commit for run:
07798b704e43036ad05cb87cdb6da91d77e5c160 / kary zheng <[email protected]>
fix(operator): read what the executor reads in the exported sources

Five places where an exported source still answered differently from the
operator the workflow ran:

- JSONL: a column the operator types STRING holds the text Jackson gives
  each value, so 35.0 is "35.0", true is "true" and a null is "null". A
  DOUBLE whose values are all whole stays a double, an INTEGER missing a
  key stays an integer, and without flattening a nested object or array
  is no column.
- CSV: a column holding a null keeps the type the schema declares, where
  pandas widened an INTEGER to a float and a BOOLEAN to an object.
- Every scan: a read of no rows takes the schema's columns in the
  schema's types. A JSONL read of no lines had no columns at all.
- CSV: the rows to skip are asked of each line. A list of two billion
  ran the script out of memory.
- Old CSV, in the operator itself: the header line and the offset are
  dropped one after the other, since their sum wraps past what an Int
  holds, and a window past the last row is inferred from the first row
  instead of leaving no types.

The spec tests that each asked one of these questions of a file of its
own now lean on the source runner's shared table, which reads every
source through the same cases and checks each column's dtype as well as
its values. Taking out any of those fixes turns it red. What it cannot
reach stays here: Parallel CSV, which the suite does not enumerate, the
schema an operator infers, which both readings share, and a failure,
which a comparison of two outputs cannot see.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

Report URL: https://github.com/apache/texera/actions/runs/36286186271

With regards,
GitHub Actions via GitBox

Reply via email to