parthchandra opened a new pull request, #6784:
URL: https://github.com/apache/datafusion-comet/pull/6784
## Which issue does this PR close?
Part of #6008.
## Rationale for this change
Comet cannot natively scan Spark's `text` file format, reporting `[COMET:
Unsupported file format Text]`.
This cascades on the build side of a broadcast join: when a small lookup
table is read from a text file,
the build branch (text scan → … → `BroadcastExchange`) stays on Spark, the
exchange never becomes a
`CometBroadcastExchange`, and the `BroadcastHashJoin` cannot run natively
even when the large probe side
is fully native. A tiny text lookup can disqualify native acceleration of
the whole join.
This PR implements option B from #6008: a native Text reader (single `value:
string` column,
line-delimited), mirroring the existing native CSV V2 scan path, so a text
leaf no longer blocks native
execution.
## What changes are included in this PR?
New config `spark.comet.scan.text.v2.enabled` (default `false`,
experimental). As with the native CSV
scan, only Spark's DataSource V2 text path is accelerated, so `text` must
also be removed from
`spark.sql.sources.useV1SourceList`.
**Native (Rust):**
- New `TextSource`/`TextOpener`
(`native/core/src/execution/operators/text_scan.rs`): a DataFusion
`FileSource`/`FileOpener` producing a single nullable `value: Utf8`
column. Line splitting matches
Spark/Hadoop `LineRecordReader`: default universal newline (`\n`, `\r\n`,
`\r`), a custom `lineSep`
(the exact charset-encoded `lineSeparatorInRead` bytes), or
`wholetext=true` (one row per file). A
leading UTF-8 BOM is stripped in line mode (matching
`skipUtfByteOrderMark`); wholetext keeps it.
Bytes are decoded the same lossy way Spark renders them on collect
(`decode_utf8_spark_lossy`). Native
text is restricted to unsplit files (`CometScanRule` falls back for split
files); a defensive
byte-range filter keeps results correct even if a split reached the
reader. Each batch's value buffer
is bounded below Arrow's 2GB offset limit, returning a clean error on a
single oversized value rather
than panicking.
- New `TextScan`/`TextOptions` protobuf messages, planner wiring
(`OpStruct::TextScan`), and
operator-registry / `jni_api` entries.
**JVM (Scala):**
- `CometTextNativeScanExec` (mirrors `CometCsvNativeScanExec`), the config,
and dispatch in
`CometExecRule`.
- `CometScanRule` claims a V2 `TextScan` only when safe, otherwise falling
back to Spark: native exec
disabled,
`input_file_name`/`input_file_block_start`/`input_file_block_length`, a
multi-column read
schema (matching Spark's `TextScan.verifyReadSchema`), partition columns,
metadata columns, compressed
files, files Spark split into byte ranges, oversized `wholetext` files,
`ignoreCorruptFiles`/
`ignoreMissingFiles`, unsupported filesystem schemes, and
object_store-rejected paths. This scopes
native text to the small, unsplit, uncompressed lookup files the feature
targets.
**Docs:** a Text subsection in the data sources guide.
## How are these changes tested?
- Rust unit tests in `text_scan.rs`: line splitting (universal newline,
trailing terminator,
empty/interior-empty lines, custom single- and multi-byte separators,
wholetext), byte-range split
partitioning, BOM stripping, invalid-UTF-8 → U+FFFD, and the
count/empty-projection path.
- `CometTextNativeReadSuite` compares Comet to Spark
(`checkSparkAnswerAndOperator`) for basic read,
count, filter, wholetext (+ count, directory), custom `lineSep` (incl. a
trailing separator and a
non-ASCII multibyte `。` separator), CRLF/CR line endings, invalid UTF-8,
BOM, multi-file directories,
and empty files. It asserts the correct fallback
(`checkSparkAnswerAndFallbackReason`) for split
files, compressed files, `ignoreCorruptFiles`/`ignoreMissingFiles`,
`input_file_name`, large
wholetext, and native-exec-disabled, and includes a within-file
line-ordering check (direct
collect-and-compare). The suite is registered in both `pr_build_linux.yml`
and `pr_build_macos.yml`.
- Compiles against the Spark 4.0 and 4.1 profiles.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]