aminghadersohi opened a new pull request, #44951:
URL: https://github.com/apache/superset/pull/44951

   ## TL;DR
   - Let callers shrink large dashboard dataset responses by limiting column 
details.
   - Preserve the existing output by default, including counts, metrics, and 
access filtering.
   - Make oversized-response advice point to the supported column cap.
   
   ### SUMMARY
   
   #### Why
   
   `get_dashboard_datasets` returns up to 100 columns for every dataset, but 
previously accepted only a dashboard identifier. Several wide datasets can 
exceed the response byte limit with no way to request a smaller result.
   
   #### What
   
   Add one optional `max_columns` request parameter, bounded to 0–100 and 
defaulting to 100. A value of 0 returns empty column lists while preserving 
total column counts, truncation indicators, dataset identity, chart counts, and 
metrics. Apply the same cap to SQL datasets and semantic views. Document the 
parameter and suggest it in response-size errors; if columns are already 
omitted, explain the remaining limitation and suggest individual dataset 
lookups instead.
   
   This deliberately bounds column details, not dataset count or total 
serialized bytes. The response can still exceed a byte limit if its remaining 
metadata is too large. No permission checks or provider-failure behavior change.
   
   ### BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF
   
   Not applicable: no UI changes.
   
   ### TESTING INSTRUCTIONS
   
   1. Call `get_dashboard_datasets` with `{"request":{"identifier":123}}` and 
confirm its output remains unchanged.
   2. Repeat with `max_columns: 1`, then `max_columns: 0`. Confirm the column 
lists shrink, total counts remain accurate, and metrics are retained.
   3. Confirm negative values and values above 100 fail request validation.
   4. Run:
   
   ```bash
   PYTHONPATH="$PWD/superset-core/src:$PWD" pytest \
     tests/unit_tests/mcp_service/dashboard/tool/test_get_dashboard_datasets.py 
\
     tests/unit_tests/mcp_service/dashboard/test_dashboard_schemas.py \
     tests/unit_tests/mcp_service/utils/test_response_size_utils.py -q
   ```
   
   #### Eval evidence
   
   - 283 targeted unit tests passed, including 12 added cases covering 
parameter bounds/defaults, discovery schema, table and semantic-view caps, 
byte-budget reduction, and tool-specific error advice.
   - The regression run before implementation had 9 failures; the corresponding 
cases pass with the fix.
   - Changed-file pre-commit hooks, including mypy, ruff, and pylint, passed.
   - No live-server acceptance or end-to-end client evaluation was performed.
   
   #### Cost & latency delta
   
   A synthetic dashboard with 10 datasets × 120 columns measured 92,183 
serialized bytes at the default cap versus 3,303 bytes at `max_columns=0` 
(96.4% smaller). Over 100 local iterations, serializer-plus-size-measurement 
p50/p95 was 5.912/7.103 ms versus 0.276/0.361 ms. This is a mocked metadata 
microbenchmark, not warehouse or end-to-end latency. No paid API calls are 
introduced; tokenizer-specific token counts and client costs were not measured.
   
   ### ADDITIONAL INFORMATION
   
   #### Risk & rollback
   
   Backward-compatible default and unchanged response schema. No migration or 
new feature flag. Revert the change to remove the optional parameter. Review 
the truncation/count behavior and the zero-cap size advice in particular.
   
   - [ ] Has associated issue:
   - [ ] Required feature flags:
   - [ ] Changes UI
   - [ ] Includes DB Migration (follow approval process in 
[SIP-59](https://github.com/apache/superset/issues/13351))
     - [ ] Migration is atomic, supports rollback & is backwards-compatible
     - [ ] Confirm DB migration upgrade and downgrade tested
     - [ ] Runtime estimates and downtime expectations provided
   - [x] Introduces new feature or API
   - [ ] Removes existing feature or API
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to