SEPURI-SAI-KRISHNA opened a new issue, #45085:
URL: https://github.com/apache/superset/issues/45085
### Bug description
`geohash_decode` and `geodetic_parse` build their parsed frame as a bare
`DataFrame()`, which gets a fresh `RangeIndex`. They then hand it to
`_append_columns`, which aligns on the index. If the caller's frame is indexed
by anything other than 0..n-1, nothing aligns: the real rows get no
coordinates, and the parsed values arrive as extra rows.
Reproduction on current master, with a plain two-operation chain and no
contrived index:
```python
import superset.utils.pandas_postprocessing as pp
from pandas import DataFrame
df = DataFrame({
"region": ["east", "east", "west", "west"],
"geodetic": ["41.1 -73.5", "41.1 -73.5", "40.7 -74.0", "40.7 -74.0"],
})
pivoted = pp.pivot(df=df, index=["region"], columns=[],
aggregates={"geodetic": {"operator": "max"}})
pp.geodetic_parse(df=pivoted, geodetic="geodetic", latitude="lat",
longitude="lon")
# geodetic lat lon
# east 41.1 -73.5 NaN NaN
# west 40.7 -74.0 NaN NaN
# 0 NaN 41.1 -73.5
# 1 NaN 40.7 -74.0
```
Two rows in, four rows out. Every real row loses its coordinates, and two
rows appear that belong to no region.
`geohash_decode` behaves the same way. `geohash_encode` is **not** affected:
it slices `df[[latitude, longitude]]`, so it inherits the caller's index.
A non-`RangeIndex` frame mid-chain is an expected state rather than an
exotic one. `flatten` already branches on it at
`superset/utils/pandas_postprocessing/flatten.py:103`:
```python
if reset_index and not isinstance(df.index, pd.RangeIndex):
df = df.reset_index(level=0)
```
Found by a review bot on #45018, where I confirmed it predates that PR. It
survives the merge, so it is filed on its own.
The index has never been carried. Before #19116, `_append_columns` was a
plain `base_df.assign(...)`, which aligns on the index too, so the same input
came back with silently null coordinates and the correct row count:
```
city lat lon
5 A NaN NaN
6 B NaN NaN
```
#19116 replaced that with `pd.concat(axis="columns")`, which unions the
index instead, turning the silent nulls into extra rows as well. Both are
wrong, and the cause in both is the bare `DataFrame()`.
### How to reproduce the bug
1. Create a chart over a dataset with a geodetic or geohash string column.
2. Add a `pivot` post-processing operation grouping by some column.
3. Add a `geodetic_parse` or `geohash_decode` operation after it.
4. The response has more rows than groups, the real rows have null
coordinates, and the coordinates sit on rows with no group label.
### Screenshots/recordings
_No response_
### Superset version
master / latest-dev
### Python version
3.11
### Node version
Not applicable
### Browser
Not applicable
### Additional context
The fix is to construct the parsed frame on the caller's index,
`DataFrame(index=df.index)`, in the two functions that build a fresh frame.
That makes the alignment `_append_columns` performs a no-op rather than a
mismatch, and it is correct whichever branch of `_append_columns` the mapping
takes, so it does not depend on #45083.
### Checklist
- [x] I have searched Superset docs and Slack and didn't find a solution to
my problem.
- [x] I have searched the GitHub issue tracker and didn't find a similar bug
report.
- [x] I have checked Superset's logs for errors and if I found a relevant
Python stacktrace, I included it here as text or in a screenshot.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]