The GitHub Actions job "Backport Approval Check" on texera.git/fix/8596-filescan-encoding has failed. Run started by GitHub user xuang7 (triggered by xuang7).
Head commit for run: b55f9176db79e5102fb7c7d325432e4688b3a6c1 / suyashj1231 <[email protected]> fix(workflow-operator): honor the File Scan Encoding field FileScanSourceOpDesc re-declares the inherited fileEncoding as its own `encoding` property so the field can carry a hide annotation, and suppresses the inherited one with @JsonIgnoreProperties(value = Array("limit", "offset", "fileEncoding")) That is the same pattern TextSourceOpDesc uses for fileScanLimit and fileScanOffset, and those two work. Encoding did not, for two reasons: `encoding` was `private val`, so the executor could not read it, and the executor instead read `desc.fileEncoding`, which the annotation strips during serialization and which therefore always came back as its UTF_8 default. So the Encoding field on a File Scan changed nothing. A UTF-16 file was decoded as UTF-8 whatever the user picked. Make `encoding` a `var`, matching ScanSourceOpDesc.fileEncoding and FileScanOpDesc.fileEncoding, and read it in the executor. FileScanOpDesc declares its own public fileEncoding and was never affected. The existing US_ASCII case set the inherited field, so it exercised the encoding path without being able to detect this: ASCII and UTF-8 agree on ASCII bytes, so it passed either way. It now sets `encoding`, and two cases are added — one that the charset survives the descriptor round trip getPhysicalOp uses to reach the executor, and one that a real UTF-16 file decodes to its text. The latter fails on the old wiring. Closes #8596 Claude-Session: https://claude.ai/code/session_01EeaEYRdhRYWL7ya7w8LJux Report URL: https://github.com/apache/texera/actions/runs/35469070062 With regards, GitHub Actions via GitBox
