[ 
https://issues.apache.org/jira/browse/CAMEL-25075?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18122052#comment-18122052
 ] 

Claus Ibsen commented on CAMEL-25075:
-------------------------------------

_Claude Code on behalf of davsclaus_

PR: https://github.com/apache/camel/pull/27298 (parts 1 and 2 of the plan 
above; part 3, the measurement, is for the next benchmark run)

* Deterministic tools: camel_catalog_doc, camel_catalog_find, 
camel_catalog_sample, camel_error_diagnose. Listed by both servers with 
"_meta": {"camel.apache.org/deterministic": true}.
* From the third identical call of a session, a short note instead of the full 
answer (checked live: 17,073 characters, then 265). No advice attached.
* Session: an MCP connection for camel mcp; from one initialize to the next for 
camel tui --mcp.
* The TUI AI panel is unchanged: it already stops identical calls within a 
turn, and it compacts older results between turns, so a session-wide count 
would cut the re-read it asks for.

One thing to check in the benchmark setup: if the harness keeps one MCP 
connection open across steps that are separate agent conversations, a new 
step's third ask of the same lookup would get the note although that 
conversation never saw the answer. One connection or initialize per step avoids 
it.

> camel-jbang-mcp - a repeated identical tool call should answer differently, 
> so an agent can notice it is looping
> ----------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25075
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25075
>             Project: Camel
>          Issue Type: Improvement
>          Components: camel-jbang
>            Reporter: Claus Ibsen
>            Assignee: Claus Ibsen
>            Priority: Major
>
> An agent that asks the same tool the same question over and over gets the 
> same answer each time, with nothing in the answer to tell it so. A local 
> model does this, and it is the largest single cause of failure on one rung of 
> the benchmark.
> Measured on the {{connect-service-sql}} example, whose first step writes a 
> route that inserts a row with named parameters. Across 28 attempts in three 
> 10-pass runs, the calls of a failing attempt look like this (s19-2, verbatim):
> {noformat}
> camel_get_files()
> camel_get_files(sql.camel.yaml)
> camel_get_files(application.properties)
> camel_get_files(orders/order-1001.json)
> camel_catalog_doc(sql)
> camel_catalog_doc(sql)
> ... 16 times in total, identical arguments
> {noformat}
> It never wrote a file, and the step failed with its whole budget spent. Of 10 
> passes of that step, 8 ended pinned at the tool-call ceiling, and in most of 
> them the majority of the calls were identical repeats of 
> {{camel_catalog_doc(sql)}}. Raising the ceiling does not help: the same step 
> was run at 12, 20 and 24 calls, and at 24 the extra calls were spent on more 
> repeats.
> What it was hunting for was the SQL dialect of the database, which the 
> catalog cannot supply, so the answer was never going to change. But a model 
> has no way of noticing that the answer it just received is the one it already 
> had.
> This is fixable outside the model. Some options, in rough order of how little 
> they assume:
> * the server notices that a call repeats an earlier call of the same session 
> with the same arguments, and answers with something different: that the 
> answer is unchanged, how many times it has been asked, and what else is worth 
> trying ({{camel_catalog_docs}} for the page, {{camel_catalog_sample}} for a 
> shape, {{camel_get_files}} for what is on disk);
> * the repeat answer carries a short list of the tools not yet used in this 
> session, since an agent going in circles is usually one that has not thought 
> of the next tool;
> * a cheaper variant: the answer is returned as before with one line prepended 
> saying it is a repeat.
> The point is not to refuse the call -- a legitimate re-read after an edit 
> must still work -- but to make a repeat *look different* so it can break the 
> loop.
> Filed from the local-model benchmark of the camel-jbang-mcp server; see 
> CAMEL-25040 for the related finding that the catalog answer omitted what the 
> author needed in the first place.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to