adityamparikh opened a new issue, #256:
URL: https://github.com/apache/solr-mcp/issues/256

   ## Problem
   The integration tests verify what each tool does once it is called. Nothing
   checks the part an MCP client actually depends on: whether the tool 
descriptions
   and parameter docs lead a model to the right tool with the right arguments, 
and
   whether its answer stays within what the tool returned. A description change 
can
   break agent behavior while every test passes. The risk grows with the tool
   surface, as overlapping tools (e.g. `check-health` and a future 
cluster-status
   tool) become harder for a model to tell apart.
   
   ## Proposal
   An opt-in eval suite that sends natural-language questions through a chat 
model
   connected to solr-mcp, records the tool calls, and judges each transcript on:
   
   - **Tool selection:** the expected tool was called. This is checked 
deterministically.
   - **Argument quality:** the arguments express the question; for `search`, 
the query and filters match the intent.
   - **Grounding:** the answer contains only what the tools returned.
   
   Start small: 5–6 cases over the existing conference dataset, aimed at likely
   confusion rather than full coverage. Include one no-match search that must 
not
   invent results.
   
   ## Intended use
   A pre-release check and a check on PRs that change tool descriptions; **not 
a CI
   gate**. Model output varies between runs, so each case runs several times and
   reports a pass rate. The suite is excluded from the default build, with a
   documented command and a line in the PR template and release checklist.
   A nightly CI job can be added later if the setup allows.
   
   ## Implementation options
   - **[Spring AI 
TypeSafe](https://github.com/spring-ai-community/spring-ai-typesafe)**
     (Apache-2.0, Spring AI 2.0.1, the same version as solr-mcp). Its `JevJudge`
     combines deterministic checks with typed judgments and reports which 
criterion
     failed and by how much. The judge can run on:
     - 
**[Laya](https://spring-ai-community.github.io/spring-ai-typesafe/latest/client/Laya/)**:
       Apache-2.0, local, no API key.
     - 
**[Ollama](https://spring-ai-community.github.io/spring-ai-typesafe/latest-snapshot/client/Ollama/)**
       0.35+: local, no API key, and can also run the agent model.
     - **Hosted Jev**: more decisive than the local models, but needs a key.
   - **A chat model as judge** through Spring AI's evaluation API. This is 
simpler to
     start with, but the verdicts are free text rather than typed scores.
   
   With Ollama or Laya the whole suite can run with no API keys.
   
   ## Acceptance criteria
   - [ ] Evals run via a dedicated command and are excluded from the default 
build
   - [ ] Runs locally with no API keys; hosted models are opt-in through 
configuration
   - [ ] Each case reports a pass rate over N runs and, on failure, which 
criterion failed
   - [ ] The README covers how to run the suite and add a case; the PR template 
and release checklist reference it
   
   ## Open questions
   1. Which local agent model is reliable enough at tool calling to make 
results meaningful?
   2. Should we add a nightly CI job (CPU-only runners, model downloads), or 
run locally only?
   3. What pass rate should count as a regression?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to