sadpandajoe commented on code in PR #43693:
URL: https://github.com/apache/superset/pull/43693#discussion_r3989461944
##########
superset/utils/pandas_postprocessing/pivot.py:
##########
@@ -187,6 +187,25 @@ def _restore_dropped_metric_columns(
return df
+def _fill_dimension_column(df: DataFrame, col: str, fill_value: str) -> None:
+ """Fill missing values in a groupby dimension column before pivoting.
+
+ Handles categorical dtypes (adding fill_value to categories) and datetime
+ dtypes (converting to string representation with fill_value for NaT) to
prevent
+ dtype errors and preserve NULL/NaN/NaT keys through pivot_table().
+ """
+ s = df[col]
+ if isinstance(s.dtype, pd.CategoricalDtype) and fill_value not in
s.cat.categories:
+ df[col] = s.cat.add_categories([fill_value]).fillna(value=fill_value)
+ elif pd.api.types.is_datetime64_any_dtype(s.dtype) or (
+ getattr(s.dtype, "kind", None) == "M"
+ ):
+ if s.isna().any():
+ df[col] = s.astype(str).where(~s.isna(), other=fill_value)
+ else:
+ df[col] = s.fillna(value=fill_value)
Review Comment:
Nullable pandas extension dimensions (`Int64`, `Float64`, or `boolean`)
reach this branch, where filling with `"<NULL>"` raises `TypeError` instead of
preserving the null group, so the pivot fails outright. Could this handle
extension dtypes before filling, with a regression using an `Int64` series
containing `pd.NA`?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]