[ https://issues.apache.org/jira/browse/SPARK-28266?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]
Erik Krogen reopened SPARK-28266: --------------------------------- Re-opening this issue based on [~shardulm]'s example above demonstrating that this is indeed a real-world issue. We have faced this issue internally on a number of occasions and have been operating a fix, on which [~shardulm]'s PR is based, for a few years. > data duplication when `path` serde property is present > ------------------------------------------------------ > > Key: SPARK-28266 > URL: https://issues.apache.org/jira/browse/SPARK-28266 > Project: Spark > Issue Type: Bug > Components: Spark Core > Affects Versions: 2.2.0, 2.2.1, 2.2.2 > Reporter: Ruslan Dautkhanov > Priority: Major > Labels: correctness > > Spark duplicates returned datasets when `path` serde is present in a parquet > table. > Confirmed versions affected: Spark 2.2, Spark 2.3, Spark 2.4. > Confirmed unaffected versions: Spark 2.1 and earlier (tested with Spark 1.6 > at least). > Reproducer: > {code:python} > >>> spark.sql("create table ruslan_test.test55 as select 1 as id") > DataFrame[] > >>> spark.table("ruslan_test.test55").explain() > == Physical Plan == > HiveTableScan [id#16], HiveTableRelation `ruslan_test`.`test55`, > org.apache.hadoop.hive.serde2.lazy.LazySimpleSerDe, [id#16] > >>> spark.table("ruslan_test.test55").count() > 1 > {code} > (all is good at this point, now exist session and run in Hive for example - ) > {code:sql} > ALTER TABLE ruslan_test.test55 SET SERDEPROPERTIES ( > 'path'='hdfs://epsdatalake/hivewarehouse/ruslan_test.db/test55' ) > {code} > So LOCATION and serde `path` property would point to the same location. > Now see count returns two records instead of one: > {code:python} > >>> spark.table("ruslan_test.test55").count() > 2 > >>> spark.table("ruslan_test.test55").explain() > == Physical Plan == > *(1) FileScan parquet ruslan_test.test55[id#9] Batched: true, Format: > Parquet, Location: > InMemoryFileIndex[hdfs://epsdatalake/hivewarehouse/ruslan_test.db/test55, > hdfs://epsdatalake/hive..., PartitionFilters: [], PushedFilters: [], > ReadSchema: struct<id:int> > >>> > {code} > Also notice that the presence of `path` serde property makes TABLE location > show up twice - > {quote} > InMemoryFileIndex[hdfs://epsdatalake/hivewarehouse/ruslan_test.db/test55, > hdfs://epsdatalake/hive..., > {quote} > We have some applications that create parquet tables in Hive with `path` > serde property > and it makes data duplicate in query results. > Hive, Impala etc and Spark version 2.1 and earlier read such tables fine, but > not Spark 2.2 and later releases. -- This message was sent by Atlassian Jira (v8.3.4#803005) --------------------------------------------------------------------- To unsubscribe, e-mail: issues-unsubscr...@spark.apache.org For additional commands, e-mail: issues-h...@spark.apache.org