adityamparikh opened a new issue, #256: URL: https://github.com/apache/solr-mcp/issues/256
## Problem The integration tests verify what each tool does once it is called. Nothing checks the part an MCP client actually depends on: whether the tool descriptions and parameter docs lead a model to the right tool with the right arguments, and whether its answer stays within what the tool returned. A description change can break agent behavior while every test passes. The risk grows with the tool surface, as overlapping tools (e.g. `check-health` and a future cluster-status tool) become harder for a model to tell apart. ## Proposal An opt-in eval suite that sends natural-language questions through a chat model connected to solr-mcp, records the tool calls, and judges each transcript on: - **Tool selection:** the expected tool was called. This is checked deterministically. - **Argument quality:** the arguments express the question; for `search`, the query and filters match the intent. - **Grounding:** the answer contains only what the tools returned. Start small: 5–6 cases over the existing conference dataset, aimed at likely confusion rather than full coverage. Include one no-match search that must not invent results. ## Intended use A pre-release check and a check on PRs that change tool descriptions; **not a CI gate**. Model output varies between runs, so each case runs several times and reports a pass rate. The suite is excluded from the default build, with a documented command and a line in the PR template and release checklist. A nightly CI job can be added later if the setup allows. ## Implementation options - **[Spring AI TypeSafe](https://github.com/spring-ai-community/spring-ai-typesafe)** (Apache-2.0, Spring AI 2.0.1, the same version as solr-mcp). Its `JevJudge` combines deterministic checks with typed judgments and reports which criterion failed and by how much. The judge can run on: - **[Laya](https://spring-ai-community.github.io/spring-ai-typesafe/latest/client/Laya/)**: Apache-2.0, local, no API key. - **[Ollama](https://spring-ai-community.github.io/spring-ai-typesafe/latest-snapshot/client/Ollama/)** 0.35+: local, no API key, and can also run the agent model. - **Hosted Jev**: more decisive than the local models, but needs a key. - **A chat model as judge** through Spring AI's evaluation API. This is simpler to start with, but the verdicts are free text rather than typed scores. With Ollama or Laya the whole suite can run with no API keys. ## Acceptance criteria - [ ] Evals run via a dedicated command and are excluded from the default build - [ ] Runs locally with no API keys; hosted models are opt-in through configuration - [ ] Each case reports a pass rate over N runs and, on failure, which criterion failed - [ ] The README covers how to run the suite and add a case; the PR template and release checklist reference it ## Open questions 1. Which local agent model is reliable enough at tool calling to make results meaningful? 2. Should we add a nightly CI job (CPU-only runners, model downloads), or run locally only? 3. What pass rate should count as a regression? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
