rusackas commented on code in PR #43693:
URL: https://github.com/apache/superset/pull/43693#discussion_r4067520295
##########
superset/utils/pandas_postprocessing/pivot.py:
##########
@@ -187,6 +187,25 @@ def _restore_dropped_metric_columns(
return df
+def _fill_dimension_column(df: DataFrame, col: str, fill_value: str) -> None:
+ """Fill missing values in a groupby dimension column before pivoting.
+
+ Handles categorical dtypes (adding fill_value to categories) and datetime
+ dtypes (converting to string representation with fill_value for NaT) to
prevent
+ dtype errors and preserve NULL/NaN/NaT keys through pivot_table().
+ """
+ s = df[col]
+ if isinstance(s.dtype, pd.CategoricalDtype) and fill_value not in
s.cat.categories:
Review Comment:
Fixed in `bf359f7`: only add the category and fill when `s.isna().any()`, so
a categorical dimension with no missing values renders exactly as before
instead of picking up an unobserved `<NULL>` bucket. Added a test asserting
`NULL_STRING` stays out of the index when there's nothing to fill.
##########
superset/utils/pandas_postprocessing/pivot.py:
##########
@@ -187,6 +187,25 @@ def _restore_dropped_metric_columns(
return df
+def _fill_dimension_column(df: DataFrame, col: str, fill_value: str) -> None:
+ """Fill missing values in a groupby dimension column before pivoting.
+
+ Handles categorical dtypes (adding fill_value to categories) and datetime
+ dtypes (converting to string representation with fill_value for NaT) to
prevent
+ dtype errors and preserve NULL/NaN/NaT keys through pivot_table().
+ """
+ s = df[col]
+ if isinstance(s.dtype, pd.CategoricalDtype) and fill_value not in
s.cat.categories:
+ df[col] = s.cat.add_categories([fill_value]).fillna(value=fill_value)
+ elif pd.api.types.is_datetime64_any_dtype(s.dtype) or (
+ getattr(s.dtype, "kind", None) == "M"
+ ):
+ if s.isna().any():
+ df[col] = s.astype(str).where(~s.isna(), other=fill_value)
+ else:
+ df[col] = s.fillna(value=fill_value)
Review Comment:
Fixed in `bf359f7`: casting the extension-dtype column to object before
filling, since a masked `Int64`/`Float64`/`boolean` array can't hold the string
sentinel directly. Added a regression with an `Int64` index carrying `pd.NA`.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]