sadpandajoe commented on code in PR #43694:
URL: https://github.com/apache/superset/pull/43694#discussion_r3890741089


##########
superset/utils/pandas_postprocessing/pivot.py:
##########
@@ -187,6 +187,31 @@ def _restore_dropped_metric_columns(
     return df
 
 
+def _fill_dimension_column(df: DataFrame, col: str, fill_value: str) -> None:
+    """Fill missing values in a groupby dimension column before pivoting.
+
+    Handles categorical dtypes (adding fill_value to categories) and datetime
+    dtypes (converting to string representation with fill_value for NaT) to 
prevent
+    dtype errors and preserve NULL/NaN/NaT keys through pivot_table().
+    """
+    s = df[col]
+    if (
+        isinstance(s.dtype, pd.CategoricalDtype)
+        and fill_value not in s.cat.categories
+    ):
+        df[col] = s.cat.add_categories([fill_value]).fillna(value=fill_value)

Review Comment:
   Agreed—adding `<NULL>` to every categorical dimension creates a zero-valued 
group when no input is null, so pivots report a group that does not exist. 
Should this only add the category when `s.isna().any()`?



##########
superset/utils/pandas_postprocessing/pivot.py:
##########
@@ -187,6 +187,31 @@ def _restore_dropped_metric_columns(
     return df
 
 
+def _fill_dimension_column(df: DataFrame, col: str, fill_value: str) -> None:
+    """Fill missing values in a groupby dimension column before pivoting.
+
+    Handles categorical dtypes (adding fill_value to categories) and datetime
+    dtypes (converting to string representation with fill_value for NaT) to 
prevent
+    dtype errors and preserve NULL/NaN/NaT keys through pivot_table().
+    """
+    s = df[col]
+    if (
+        isinstance(s.dtype, pd.CategoricalDtype)
+        and fill_value not in s.cat.categories
+    ):
+        df[col] = s.cat.add_categories([fill_value]).fillna(value=fill_value)
+    elif pd.api.types.is_datetime64_any_dtype(s.dtype) or getattr(s.dtype, 
"kind", None) == "M":
+        if s.isna().any():
+            df[col] = s.astype(str).replace({
+                "NaT": fill_value,
+                "<NA>": fill_value,
+                "nan": fill_value,
+                "None": fill_value,
+            })
+    else:
+        df[col] = s.fillna(value=fill_value)

Review Comment:
   Agreed—turning non-null timestamps into strings changes the pivot labels, 
and interval dimensions with a null still fail at `fillna`. Could this preserve 
temporal values as objects while replacing only missing entries?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to