The GitHub Actions job "Backport Approval Check" on 
texera.git/fix/8596-filescan-encoding has failed.
Run started by GitHub user xuang7 (triggered by xuang7).

Head commit for run:
b55f9176db79e5102fb7c7d325432e4688b3a6c1 / suyashj1231 
<[email protected]>
fix(workflow-operator): honor the File Scan Encoding field

FileScanSourceOpDesc re-declares the inherited fileEncoding as its own
`encoding` property so the field can carry a hide annotation, and
suppresses the inherited one with

  @JsonIgnoreProperties(value = Array("limit", "offset", "fileEncoding"))

That is the same pattern TextSourceOpDesc uses for fileScanLimit and
fileScanOffset, and those two work. Encoding did not, for two reasons:
`encoding` was `private val`, so the executor could not read it, and the
executor instead read `desc.fileEncoding`, which the annotation strips
during serialization and which therefore always came back as its
UTF_8 default.

So the Encoding field on a File Scan changed nothing. A UTF-16 file was
decoded as UTF-8 whatever the user picked.

Make `encoding` a `var`, matching ScanSourceOpDesc.fileEncoding and
FileScanOpDesc.fileEncoding, and read it in the executor. FileScanOpDesc
declares its own public fileEncoding and was never affected.

The existing US_ASCII case set the inherited field, so it exercised the
encoding path without being able to detect this: ASCII and UTF-8 agree
on ASCII bytes, so it passed either way. It now sets `encoding`, and two
cases are added — one that the charset survives the descriptor round
trip getPhysicalOp uses to reach the executor, and one that a real
UTF-16 file decodes to its text. The latter fails on the old wiring.

Closes #8596

Claude-Session: https://claude.ai/code/session_01EeaEYRdhRYWL7ya7w8LJux

Report URL: https://github.com/apache/texera/actions/runs/35469070062

With regards,
GitHub Actions via GitBox

Reply via email to