I3eka opened a new pull request, #44717:
URL: https://github.com/apache/superset/pull/44717

   ### SUMMARY
   
   Give browser-supplied page context the same untrusted-content framing 
already used for native AI tool data. This addresses the [page-context review 
finding in 
#43237](https://github.com/apache/superset/pull/43237#discussion_r4002625804), 
also called out in the maintainer's 
[follow-up](https://github.com/apache/superset/pull/43237#issuecomment-5801835163).
   
   - Reuse `sanitize_for_llm_context` for rendered titles, paths, SQL, 
Markdown, table names, chart controls and filter values. Embedded framing 
delimiters are escaped by the existing helper.
   - Leave allowlisted page types and validated numeric object IDs outside the 
data frames, so the assistant can still use the existing tool arguments and 
guidance.
   - Apply per-field limits before framing and cap the overall section at 
complete rendered entries. A size limit cannot leave an open data frame or an 
incomplete SQL block.
   
   Both the conversation orchestrator and opening suggestions already call this 
renderer; no per-caller sanitizer, dependency or alternate prompt assembly path 
is added. Framing is guidance to the model, not a replacement for tool 
authorization or proof of resistance to every prompt injection. This is a 
consistency/hardening fix for the review request, not a new vulnerability 
report.
   
   **Dependencies and scope:** depends on the unmerged AI base #42805. The 
focused change is [three files above shared compatibility base 
`e809854983f396f4824d682935d1300f751d7164`](https://github.com/I3eka/superset/compare/e809854983f396f4824d682935d1300f751d7164...fix-ai-page-context-framing),
 at `07fbd7aae05725519637dce47d9204413c8f915b`. The Apache-master comparison 
includes the prerequisite. This draft does not bypass its SIP/review gate or 
include the separate idempotency/profile/model fixes.
   
   Not included: structured request validation, stored-context/replay 
semantics, profile selection, model pinning, queue recovery, retention, or 
migration reconciliation on other branches.
   
   ### BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF
   
   Backend prompt rendering only; no visual redesign.
   
   Before, arbitrary page text is interpolated into the system prompt after a 
prose warning. After, each rendered text value is visibly framed as data and 
embedded delimiters are escaped; operational IDs remain plain. Whole-context 
truncation drops entries that no longer fit instead of cutting their closing 
delimiter.
   
   ### TESTING INSTRUCTIONS
   
   - Five added cases fail on the unchanged shared base: one checks every 
rendered client-text field and retained operational IDs, and four check SQL, 
Markdown, compact-value and whole-context limits.
   - **73 focused checks pass**, covering page context and its surrounding 
tool-display, suggestions and orchestrator suites.
   - **708 backend checks pass:** the full AI unit suite, the real 
single-migration-head test and seven translation-template checks. The first 
translation invocation lacked `pybabel` on PATH; the corrected full run uses 
the existing backend virtual environment. No product workaround was added.
   - All applicable pre-commit hooks pass over **115 PR files**, including 
mypy, frontend type checking, Ruff, pylint and formatting. The existing 
referenced TypeScript declarations were rebuilt after branch switching, and 
formatting uses the locked oxfmt 0.68.0.
   - These are deterministic local rendering/contract tests, not an 
external-model injection benchmark, a live broker test or a deployed-browser 
check.
   
   ```bash
   pytest -q tests/unit_tests/ai 
tests/unit_tests/migrations/test_single_migration_head.py 
tests/unit_tests/scripts/translations/check_pot_drift_test.py
   ```
   
   To inspect manually, render a dashboard context containing instruction-like 
text and literal framing delimiters in its title, notes and filter values, plus 
editor SQL. Check that the text is inside balanced untrusted frames, embedded 
delimiters are escaped, numeric IDs remain usable, and the whole result remains 
within `MAX_CONTEXT_CHARS` even for oversized input.
   
   No deployment, configuration change, service restart, database migration or 
warehouse query accompanies this fix. Remote CI remains separate from local 
checks.
   
   ### ADDITIONAL INFORMATION
   
   - [x] Has associated issue: review finding linked above
   - [x] Required feature flags: `AI_ASSISTANT` from #42805
   - [ ] Changes UI
   - [x] Includes DB Migration (follow approval process in 
[SIP-59](https://github.com/apache/superset/issues/13351)): inherited AI 
schema/shared no-op joins only; this three-file fix adds none
     - [ ] Migration is atomic, supports rollback & is backwards-compatible: 
full inherited chain not verified by this fix
     - [ ] Confirm DB migration upgrade and downgrade tested
     - [x] Runtime estimates and downtime expectations provided: no schema/data 
operations added by this fix
   - [ ] Introduces new feature or API
   - [ ] Removes existing feature or API
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to