parthchandra opened a new pull request, #6784:
URL: https://github.com/apache/datafusion-comet/pull/6784

   ## Which issue does this PR close?
   
   Part of  #6008.
   
   ## Rationale for this change
   
   Comet cannot natively scan Spark's `text` file format, reporting `[COMET: 
Unsupported file format Text]`.
   This cascades on the build side of a broadcast join: when a small lookup 
table is read from a text file,
   the build branch (text scan → … → `BroadcastExchange`) stays on Spark, the 
exchange never becomes a
   `CometBroadcastExchange`, and the `BroadcastHashJoin` cannot run natively 
even when the large probe side
   is fully native. A tiny text lookup can disqualify native acceleration of 
the whole join.
   
   This PR implements option B from #6008: a native Text reader (single `value: 
string` column,
   line-delimited), mirroring the existing native CSV V2 scan path, so a text 
leaf no longer blocks native
   execution.
   
   ## What changes are included in this PR?
   
   New config `spark.comet.scan.text.v2.enabled` (default `false`, 
experimental). As with the native CSV
   scan, only Spark's DataSource V2 text path is accelerated, so `text` must 
also be removed from
   `spark.sql.sources.useV1SourceList`.
   
   **Native (Rust):**
   - New `TextSource`/`TextOpener` 
(`native/core/src/execution/operators/text_scan.rs`): a DataFusion
     `FileSource`/`FileOpener` producing a single nullable `value: Utf8` 
column. Line splitting matches
     Spark/Hadoop `LineRecordReader`: default universal newline (`\n`, `\r\n`, 
`\r`), a custom `lineSep`
     (the exact charset-encoded `lineSeparatorInRead` bytes), or 
`wholetext=true` (one row per file). A
     leading UTF-8 BOM is stripped in line mode (matching 
`skipUtfByteOrderMark`); wholetext keeps it.
     Bytes are decoded the same lossy way Spark renders them on collect 
(`decode_utf8_spark_lossy`). Native
     text is restricted to unsplit files (`CometScanRule` falls back for split 
files); a defensive
     byte-range filter keeps results correct even if a split reached the 
reader. Each batch's value buffer
     is bounded below Arrow's 2GB offset limit, returning a clean error on a 
single oversized value rather
     than panicking.
   - New `TextScan`/`TextOptions` protobuf messages, planner wiring 
(`OpStruct::TextScan`), and
     operator-registry / `jni_api` entries.
   
   **JVM (Scala):**
   - `CometTextNativeScanExec` (mirrors `CometCsvNativeScanExec`), the config, 
and dispatch in
     `CometExecRule`.
   - `CometScanRule` claims a V2 `TextScan` only when safe, otherwise falling 
back to Spark: native exec
     disabled, 
`input_file_name`/`input_file_block_start`/`input_file_block_length`, a 
multi-column read
     schema (matching Spark's `TextScan.verifyReadSchema`), partition columns, 
metadata columns, compressed
     files, files Spark split into byte ranges, oversized `wholetext` files, 
`ignoreCorruptFiles`/
     `ignoreMissingFiles`, unsupported filesystem schemes, and 
object_store-rejected paths. This scopes
     native text to the small, unsplit, uncompressed lookup files the feature 
targets.
   
   **Docs:** a Text subsection in the data sources guide.
   
   ## How are these changes tested?
   
   - Rust unit tests in `text_scan.rs`: line splitting (universal newline, 
trailing terminator,
     empty/interior-empty lines, custom single- and multi-byte separators, 
wholetext), byte-range split
     partitioning, BOM stripping, invalid-UTF-8 → U+FFFD, and the 
count/empty-projection path.
   - `CometTextNativeReadSuite` compares Comet to Spark 
(`checkSparkAnswerAndOperator`) for basic read,
     count, filter, wholetext (+ count, directory), custom `lineSep` (incl. a 
trailing separator and a
     non-ASCII multibyte `。` separator), CRLF/CR line endings, invalid UTF-8, 
BOM, multi-file directories,
     and empty files. It asserts the correct fallback 
(`checkSparkAnswerAndFallbackReason`) for split
     files, compressed files, `ignoreCorruptFiles`/`ignoreMissingFiles`, 
`input_file_name`, large
     wholetext, and native-exec-disabled, and includes a within-file 
line-ordering check (direct
     collect-and-compare). The suite is registered in both `pr_build_linux.yml` 
and `pr_build_macos.yml`.
   - Compiles against the Spark 4.0 and 4.1 profiles.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to