goutamadwant opened a new issue, #12240:
URL: https://github.com/apache/seatunnel/issues/12240

   ### Search before asking
   
   - [x] I searched existing issues and pull requests. This is a focused 
follow-up to the benchmark direction in #11616, not a duplicate of the existing 
comparison implementation.
   
   ### Description
   
   Add an opt-in public paraphrase regression suite to the CLI benchmark while 
keeping the existing 100-task benchmark unchanged.
   
   The current baseline uses fixed prompt wording. A separate set of equivalent 
requests would let contributors check whether a knowledge-pack or generation 
change remains stable when the same task is expressed differently. Each 
alternative prompt should reuse its original task's assertions and execution 
fixtures, so wording changes do not silently change the evaluation contract.
   
   This proposal adds evaluation infrastructure. It does not claim that a 
particular model currently fails these prompts or that model accuracy has 
improved.
   
   #### Reproduction of the missing capability
   
   Baseline: `apache/seatunnel` revision 
`75fd4ed4b2e63579a57b461285997951490c4fe7`.
   
   From `seatunnel-cli`, run:
   
   ```bash
   python -m benchmark.runner --suite paraphrase --level l1
   ```
   
   On the original baseline, this was confirmed on Python 3.10 with exit status 
2:
   
   ```text
   runner.py: error: unrecognized arguments: --suite paraphrase
   ```
   
   Argument parsing stopped before any provider call. The original loader 
returned the existing 100 tasks; there was no selectable paraphrase suite. This 
is a reproducible capability gap, not a model-quality failure.
   
   #### Proposed behavior and acceptance criteria
   
   - Add `--suite paraphrase`; retain `--suite baseline` as the default.
   - Start with 12 reviewed alternative prompts: four routing tasks, four CDC 
tasks and four connector-option/mode tasks, including one Chinese prompt. 
Selecting this suite runs the variants only, not a combined 112-task baseline.
   - Define variants using only parent task ID, full parent-contract 
fingerprint and alternative prompt text. Copy assertions, schemas, execution 
fixtures and other evaluation fields from the canonical task without allowing 
overrides.
   - Give variants distinct IDs and retain parent provenance in saved results. 
A changed parent contract requires explicit review and repinning.
   - Reject missing/duplicate parents, invalid definitions, unchanged/empty 
wording and invalid variant selections before provider setup.
   - Reuse existing generation, scoring, repair, reporting and revision 
comparison. Preserve baseline task order, definitions, fingerprints and report 
formats.
   - Document the suite and its limitations in English, Chinese and the 
benchmark README.
   
   #### Before and after
   
   Before: contributors can compare results for fixed baseline prompts, but 
cannot select a reviewed alternative-wording suite through the benchmark runner.
   
   After: contributors can run a separate public regression suite and inspect 
task-level transitions using the existing comparison tools, without replacing 
or expanding the default benchmark.
   
   #### Prepared implementation and validation
   
   Prepared commit: `8e364acfb2638ce7ca3c7c51ae035176d9e9e3ae`.
   
   From `seatunnel-cli`:
   
   ```bash
   python -m pytest tests -q
   python -m black --check benchmark/paraphrases.py 
tests/test_benchmark_paraphrases.py
   python -m ruff check --isolated --select E4,E7,E9,F benchmark/paraphrases.py 
benchmark/runner.py tests/test_benchmark_paraphrases.py
   ```
   
   - Complete CLI suite: 181 tests and 3 subtests passed on each of Python 
3.10.20 and Python 3.11.15, compared with 129 tests and 3 subtests on the 
original baseline.
   - The 52 added cases cover inheritance, mutation isolation, parent drift, 
malformed definitions, suite/tier/task selection, scoring parity, saved 
provenance, skipped execution gates and cross-revision comparison.
   - All 12 prompts were independently checked against their canonical 
requirements and fixtures. The original 100-task definitions remain identical.
   - Root formatting and the full Java 11 no-tests build passed for the branch 
and unchanged baseline. [Fork 
Build](https://github.com/goutamadwant/seatunnel/actions/runs/34320639314) also 
passed for the prepared commit.
   
   This is Python benchmark tooling, so Python 3.10/3.11 are the directly 
tested runtimes. No Java connector/runtime code is changed; Java 8 benchmark 
execution is not applicable and is not claimed. The full Maven 
compilation/package check used Java 11.
   
   The corpus is public, not an unseen holdout. Running the actual benchmark 
still calls the configured model, including L1-only runs. The reported offline 
tests made no model calls and ran no engine jobs; they do not establish 
accuracy gains, generalization, output-data equivalence or downstream 
production impact.
   
   ### Usage Scenario
   
   Regression review for CLI knowledge packs and generation changes: evaluate 
alternative wording against the same assertions, report pass-to-fail and 
fail-to-pass transitions, and preserve a reproducible default baseline for 
historical comparisons.
   
   ### Related issues
   
   - Benchmark discipline and Stage 1 direction: #11616.
   - Existing saved-result comparison: #12189 / #12192. This proposal reuses it 
rather than creating another comparator.
   - [Prepared compare 
branch](https://github.com/apache/seatunnel/compare/dev...goutamadwant:test/benchmark-paraphrase-suite).
   
   ### Are you willing to submit a PR?
   
   - [x] Yes, I am willing to submit a PR. The implementation is prepared on 
the linked compare branch.
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to