danny0405 commented on code in PR #19703:
URL: https://github.com/apache/hudi/pull/19703#discussion_r3849532662


##########
hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/spark/sql/hudi/HoodieSqlCommonUtils.scala:
##########
@@ -406,24 +406,55 @@ object HoodieSqlCommonUtils extends SparkAdapterSupport {
   private def makePartitionPath(partitionFields: Seq[String],
                                 normalizedSpecs: Map[String, String],
                                 enableEncodeUrl: Boolean,
-                                enableHiveStylePartitioning: Boolean): String 
= {
+                                enableHiveStylePartitioning: Boolean,
+                                slashSeparatedDatePartitioning: Boolean): 
String = {
+    // NOTE: Slash-separated date partitioning only kicks in for a table 
partitioned by a single
+    //       (date) column, mirroring the guard in 
[[KeyGenUtils#getRecordPartitionPath]] that drives
+    //       the write path -- these commands have to name the very same 
directory the writer created.
+    //       Hive-style partitioning is excluded because the config documents 
the two as mutually
+    //       exclusive, and the write paths do not agree on what the 
combination should produce
+    //       (tracked in HUDI issue #19669), so there is no single directory 
to name here
+    val useSlashSeparatedDates =
+      slashSeparatedDatePartitioning && !enableHiveStylePartitioning && 
partitionFields.length == 1
     partitionFields.map { partitionColumn =>
       val encodedPartitionValue = if (enableEncodeUrl) {
         
PartitionPathEncodeUtils.escapePathName(normalizedSpecs(partitionColumn))
       } else {
         normalizedSpecs(partitionColumn)
       }
-      if (enableHiveStylePartitioning) 
s"$partitionColumn=$encodedPartitionValue" else encodedPartitionValue
+      if (enableHiveStylePartitioning) {
+        s"$partitionColumn=$encodedPartitionValue"
+      } else if (useSlashSeparatedDates) {
+        toSlashSeparatedDate(encodedPartitionValue)
+      } else {
+        encodedPartitionValue
+      }
     }.mkString("/")
   }
 
+  /**
+   * Turns a `yyyy-MM-dd` formatted partition value into the `yyyy/MM/dd` 
directory structure
+   * requested by `hoodie.datasource.write.slash.separated.date.partitioning`, 
mirroring the
+   * substitution the write path performs in 
`KeyGenUtils#getRecordPartitionPath`.
+   *
+   * A value with a leading dash is returned as-is: substituting would make 
the partition path start
+   * with "/", and an absolute relative-partition-path is resolved 
inconsistently -- the writer and
+   * the file-system view disagree on where such a partition lives. Such a 
value is not a date to
+   * begin with, so nothing is lost by not slashing it.
+   */
+  private def toSlashSeparatedDate(partitionValue: String): String = {
+    if (partitionValue.startsWith("-")) partitionValue else 
partitionValue.replace('-', '/')
+  }

Review Comment:
   [P1] Keep this guard aligned with the now-merged writer path
   
   Since #19648 merged in `739ab769`, the writer suppresses substitution 
whenever `KeyGenUtils.hasPathBreakingDash` is true—not only for a leading dash, 
but also for trailing/doubled dashes and `.`/`..` tokens. This helper still 
maps values such as `2026-01-05-` to `2026/01/05/` and `2026--01` to 
`2026//01`, while the writer now stores the original dashed value. For those 
inputs, ADD again creates/checks the wrong directory and DROP/TRUNCATE silently 
target the wrong path. After rebasing, could this reuse 
`KeyGenUtils.hasPathBreakingDash(partitionValue)` (or mirror its full 
predicate), with regression coverage for at least a trailing or doubled dash?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to