SEZ9 opened a new pull request, #12604:
URL: https://github.com/apache/seatunnel/pull/12604

   ## Benchmark first
   
   `benchmark/runner.py --provider bedrock --model 
us.anthropic.claude-haiku-4-5-20251001-v1:0 --preset smoke --level l1` (12 
tasks, haiku-4-5):
   
   | | misrouted to CHAT | pass@1 | tier 3 pass@1 |
   |---|---|---|---|
   | before | **3 / 12** | 50.0% | 50.0% |
   | after | **0 / 12** | **75.0%** | **75.0%** |
   
   The three misrouted tasks are exactly the three named in #12539 
(`t1_fake_console`, `t1_mysql_localfile_csv`, `t2_mt_table_list`), each failing 
with `non-config result: chat` — the planner answered conversationally instead 
of producing a config.
   
   A 5-trial run on `t3_cr_amount_split`, the task most sensitive to plan 
shape, is **5/5 before and 5/5 after**, i.e. the change does not perturb 
multi-pipeline fan-out.
   
   Two task failures remain after the change — `t1_pg_doris` and 
`t2_mt_table_list` — both static-validation (`l1`) failures with a config 
produced, unrelated to routing.
   
   ## Purpose of this pull request
   
   Closes #12539.
   
   `PLANNER_SYSTEM`'s classification section read:
   
   > Output **PLAN:** ONLY when the user explicitly asks to CREATE or MODIFY a 
data pipeline config.
   > Signals: "sync X to Y", "read from X write to Y", "create a job that...", 
...
   >
   > Output **CHAT:** for EVERYTHING else
   
   `ONLY when` plus six literal signals plus `EVERYTHING else` reads as a 
closed verb list. Requests that plainly describe a pipeline but use a different 
verb fell through to CHAT:
   
   - `Export the inventory table from Oracle (host db1, port 1521, service 
ORCL) to HDFS as ORC files`
   - `Print 10 rows of fake data to the console`
   - `把 Oracle 的 employees 表导出到 StarRocks`
   
   The user gets prose where they asked for a config.
   
   **The parser was ruled out before the prompt was blamed.** `agents.py` 
treats a planner response as chat only on a strict `startswith("CHAT:")`, and 
unprefixed text defaults to PLAN. So the model genuinely emitted the `CHAT:` 
label; this is a classification problem, not a parsing one.
   
   ## How
   
   The section is restated as a single question — *does the request describe 
data that should end up somewhere, or a job to build?* — followed by:
   
   - verbs in English and Chinese marked **illustrative, never exhaustive**, so 
that an unlisted verb is not read as a signal to choose CHAT;
   - an explicit note that a request with no recognisable verb at all (`one 
batch job with this DAG: ...`, `三条流水线:...`) is still PLAN;
   - the unchanged CHAT list (errors, failure analysis, config review, 
connector questions, troubleshooting);
   - two **ordered** tie-breakers: pasted logs/stack traces → CHAT even when 
they name a source and a sink; otherwise a request naming data to move → PLAN 
even when it mentions testing, debugging, printing, fake data or the Console 
sink.
   
   Prompt text only — one function-free constant. No few-shot examples were 
added: an earlier variant of this change that also added three PLAN examples 
scored the same 75.0% pass@1 but introduced config-generation errors on tier-3 
fan-out tasks (`t3_cr_amount_split` 2/5 failures, `t3_ms_kafka_jdbc_es`), so 
the rule rewrite alone is what ships.
   
   ## Does this PR introduce _any_ user-facing change?
   
   No new options or commands. The planner now produces a config for pipeline 
requests that previously got a chat reply.
   
   ## Check list
   
   * [x] Code changed are covered with tests — `tests/test_planner_prompt.py`, 
195 passed
   * [x] If any new Jar binary package adding in your PR, please add License 
Notice according [New License 
Guide](https://github.com/apache/seatunnel/blob/dev/docs/en/contribution/new-license.md)
 — n/a
   * [x] If necessary, please update the documentation to describe the new 
feature — n/a, internal prompt
   * [x] If you are contributing the connector code, please check that the 
following files are updated — n/a
   
   ## Limits of the verification
   
   The benchmark suite contains only pipeline-generation tasks. It can show 
that PLAN requests stop being misrouted, but it **cannot** detect the opposite 
regression — a CHAT request newly misrouted to PLAN. The added test covers that 
direction statically: it asserts the classification section still routes stack 
traces to CHAT and is not restated as a closed verb list. It fails against the 
previous prompt.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to