SEPURI-SAI-KRISHNA commented on code in PR #19703:
URL: https://github.com/apache/hudi/pull/19703#discussion_r3835814011
##########
hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/spark/sql/hudi/HoodieSqlCommonUtils.scala:
##########
@@ -406,24 +406,55 @@ object HoodieSqlCommonUtils extends SparkAdapterSupport {
private def makePartitionPath(partitionFields: Seq[String],
normalizedSpecs: Map[String, String],
enableEncodeUrl: Boolean,
- enableHiveStylePartitioning: Boolean): String
= {
+ enableHiveStylePartitioning: Boolean,
+ slashSeparatedDatePartitioning: Boolean):
String = {
+ // NOTE: Slash-separated date partitioning only kicks in for a table
partitioned by a single
+ // (date) column, mirroring the guard in
[[KeyGenUtils#getRecordPartitionPath]] that drives
+ // the write path -- these commands have to name the very same
directory the writer created.
+ // Hive-style partitioning is excluded because the config documents
the two as mutually
+ // exclusive, and the write paths do not agree on what the
combination should produce
+ // (tracked in HUDI issue #19669), so there is no single directory
to name here
+ val slashSeparateDates =
+ slashSeparatedDatePartitioning && !enableHiveStylePartitioning &&
partitionFields.length == 1
partitionFields.map { partitionColumn =>
val encodedPartitionValue = if (enableEncodeUrl) {
PartitionPathEncodeUtils.escapePathName(normalizedSpecs(partitionColumn))
} else {
normalizedSpecs(partitionColumn)
}
- if (enableHiveStylePartitioning)
s"$partitionColumn=$encodedPartitionValue" else encodedPartitionValue
+ if (enableHiveStylePartitioning) {
+ s"$partitionColumn=$encodedPartitionValue"
+ } else if (slashSeparateDates) {
+ slashSeparateDateValue(encodedPartitionValue)
Review Comment:
Correct about master, and that mismatch is the bug #19648 fixes rather than
something this PR introduces.
On master `PartitionPathFormatterBase#combine` L64-66 is:
```java
if (slashSeparatedDatePartitioning) {
return ((S) ((String) toString(partitionPathParts[0])).replace('-', '/'));
}
```
which skips `handleEmpty` and `tryEncode` both — so for a single-field slash
table the row-writer drops URL encoding *and* null handling. #19648 rewrites
that branch to `tryEncode(handleEmpty(...))` followed by the substitution, i.e.
the same encode-then-substitute order as `KeyGenUtils#getRecordPartitionPath`
(L252-261).
So this DDL matches the Avro writer today, and the row-writer once #19648
lands. Matching master's unencoded row-writer output instead would mean
encoding that bug into the DDL and reverting it when #19648 merges.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]