dwsmith1983 opened a new issue, #6746: URL: https://github.com/apache/datafusion-comet/issues/6746
## Describe the bug The native Parquet scan builds one object store per file partition, from that partition's first file ([planner.rs#L1950-L1967](https://github.com/apache/datafusion-comet/blob/e038cf05fa665dc61fdecbdde1c07b8b6a1ce3cb/native/core/src/execution/planner.rs#L1950-L1967)), and `get_partitioned_files` keeps only each file's path, dropping the scheme and authority ([planner.rs#L491-L512](https://github.com/apache/datafusion-comet/blob/e038cf05fa665dc61fdecbdde1c07b8b6a1ce3cb/native/core/src/execution/planner.rs#L491-L512)). Spark's `FilePartition.getFilePartitions` packs small files from every root path into the same partitions, so a scan over two buckets can put files from both into one partition. The files from the second bucket are then fetched from the first. When the same key exists in both buckets the scan returns the first bucket's rows twice, with no error. `CometScanRule` already falls back for this when the paths use an opt-in S3-compatible alias scheme, and its comment notes that plain `s3://` and `s3a://` have the same flaw ([CometScanRule.scala#L293-L307](https://github.com/apache/datafusion-comet/blob/e038cf05fa665dc61fdecbdde1c07b8b6a1ce3cb/spark/src/main/scala/org/apache/comet/rules/CometScanRule.scala#L293-L307)). Plain multi-bucket scans still run natively. #6059 declines the ABFS case (more than one container or account) for the same reason. ## Steps to reproduce MinIO through `CometS3TestBase`, two buckets `repro-a` and `repro-b`. Each holds two small Parquet files written by Spark with Comet off: ids 1-3 and 4-6, plus a column `bucket` = `'A'` or `'B'`. To get all four files into one partition, set `spark.sql.files.minPartitionNum=1`, `spark.sql.files.openCostInBytes=1` and `spark.sql.files.maxPartitionBytes=128MB`. Then run: ```scala spark.read.parquet("s3a://repro-a/pq-same", "s3a://repro-b/pq-same").collect() ``` The plan is `CometNativeScan parquet ... InMemoryFileIndex(2 paths)`. Rows below are sorted after collecting. - **Same keys in both buckets** (`pq-same/f1.parquet`, `pq-same/f2.parquet`): Spark returns `[1,A] [1,B] [2,A] [2,B] ... [6,A] [6,B]`. Comet returns `[1,A] [1,A] [2,A] [2,A] ... [6,A] [6,A]`, with no error. - **Different keys** (`repro-a/pq-a/a1.parquet`, `repro-b/pq-b/b1.parquet`, ...): the scan fails with `[FAILED_READ_FILE.FILE_NOT_EXIST] ... reading file pq-b/b1.parquet`. The message carries the key without the bucket, so it does not show that the wrong bucket was queried. - **One file per partition** (`openCostInBytes=128MB`): the results match Spark. ## Expected behavior The same rows as Spark, or a fallback to Spark when the native scan cannot serve more than one bucket. ## Additional context Reproduced on `main` at e038cf05f with the default Spark 4.1 profile. The CSV native scan has a related but broader problem, filed separately. There are two ways out. The scan could decline when its files span more than one bucket or authority, which extends the alias check to every scheme. Or native planning could keep each file's store, either by grouping a partition's files by store or by carrying the full URL into the file's location. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
