[
https://issues.apache.org/jira/browse/CAMEL-25075?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18120024#comment-18120024
]
Claus Ibsen commented on CAMEL-25075:
-------------------------------------
h2. Plan
h3. What the measurements say
Across the 720 step-runs of one 10-pass benchmark run: 37% of step-runs contain
at least one identical repeated call, 556 of 5357 calls (10%) are identical
repeats. But most repeats are legitimate and must stay that way:
|| tool || identical repeats || legitimate? ||
| camel_get_files | 263 | yes -- re-reading a file after editing it |
| camel_catalog_doc | 102 | *no* -- a pure function of its arguments |
| camel_get_log | 87 | yes -- polling is the point |
| camel_write_file / camel_edit_file | 62 | mostly a retry after a refused
write |
| camel_get_errors | 21 | yes -- polling |
So the pathological share is about 2% of calls, not 10%, and it is concentrated
in the catalog tools. Worst single steps: 16x {{camel_catalog_doc}} on
route-aggregator step 2, and 14x / 11x / 9x on connect-service-sql step 1.
h3. It is a fixed point, not a memory failure
The obvious explanation -- the model can no longer see that it already asked --
is wrong. From the trace of a looping step ({{num_ctx}} 32768):
{noformat}
turn 9 prompt_eval=5278 eval=52 [camel_catalog_doc]
turn 11 prompt_eval=6547 eval=71 [camel_catalog_doc]
turn 29 prompt_eval=15831 eval=71 [camel_catalog_doc]
turn 41 prompt_eval=24063 eval=71 [camel_catalog_doc]
{noformat}
It reaches 24k of a 32k window, so nothing was evicted: every one of those
identical answers is still in context. And {{eval=71}} twelve turns running
means it emits a byte-identical response each time. The only thing changing in
the input is another copy of the same block, which perturbs nothing.
That matters for the design: *a small addition to a large identical answer is
unlikely to break it.* A repeat counter appended to 1372 identical tokens is
exactly the sort of perturbation a stuck sampler steps over. The answer has to
look different.
h3. Why the MCP hint does not fit
All 81 tools already carry annotations, set declaratively in
{{camel-jbang-mcp}} with {{@Tool(annotations = @Tool.Annotations(readOnlyHint =
...))}}: 69 read-only, 8 writers, 4 destructive. {{idempotentHint}} is false on
all 81.
Neither existing hint expresses what is needed:
* {{idempotentHint}} is defined as "repeated calls with the same arguments have
no additional effect on its environment", and the spec says it is meaningful
only when {{readOnlyHint == false}}. It is about side effects, not about the
response being identical.
* {{readOnlyHint}} cannot stand in for it. {{camel_get_log}},
{{camel_get_errors}}, {{camel_get_files}} and {{camel_catalog_doc}} are all
read-only, and only the last is deterministic.
MCP has no standard hint for "same input, same output".
h3. The plan, in three independently mergeable parts
*1. Declare it (the contract).* Add a {{deterministic}} flag to the shared tool
definition in {{camel-jbang-core}}'s tool registry, so the camel-jbang views
and the camel-jbang-mcp server agree on one source of truth, and surface it in
the MCP tool listing as a Camel-namespaced annotation (the annotations object
is open; alternatively {{_meta}}). True for the catalog and documentation
lookups, false for anything that reads the running integration, the file system
or the clock. A client can then suppress or short-circuit repeats itself,
without the server ever editing its answers -- which removes the two real
objections to part 2, the risk to the response contract and the server
second-guessing the agent. {{AuthoringToolsTest}} already asserts
{{readOnlyHint}} against the definition, so the same test gains a case.
While there: set {{idempotentHint}} where it is genuinely true among the 12
writers ({{camel_write_file}} and {{camel_render_route_diagram}} are;
{{camel_runtime_send}}, {{camel_runtime_sql}} and {{camel_runtime_heap_dump}}
are not). Minor correctness, not the mechanism.
*2. Short-circuit it (the backstop), for clients that do not read the
annotation.* From the third identical call of a deterministic tool within a
session, answer with a short payload instead of the full one: that it is
identical to the call at turn N, the repeat count, and nothing else. Two
effects, independent of each other:
* the input changes materially, which is what a fixed point needs;
* it stops spending about 1372 tokens per useless call. Those 16 calls filled
roughly 22k of context with the same text, so even if the loop survives, the
context is given back.
Deliberately no advice attached ("try camel_catalog_docs instead") in the first
cut: a wrong nudge sends the model somewhere worse, and it would make the
benchmark measure the nudge rather than the model. Add it only if the
measurement asks for it. The short answer names the turn that carried the full
one, so a legitimate re-read is not left empty-handed.
*3. Measure it.* The two worst rungs, connect-service-sql and route-aggregator,
10 passes each before and after -- about an hour a rung. What to look for: the
pinned-at-the-ceiling share (8 of 10 on connect-service-sql step 1 today), the
number of identical catalog repeats, and the step pass rate.
h3. What would make us stop
If part 2 does not reduce the repeats, do not extend it to more tools and do
not add advice: it would then be a mechanism that changes answers without
earning it, and the context saving alone is not worth that. Keep part 1 either
way, since an honest contract is worth having on its own.
And this is a second-order fix. CAMEL-25040 and CAMEL-25074 -- making an answer
carry the consequence of what it describes -- attack the cause. The
route-aggregator step 2 case (16x) is worth reading first: it may be another
answer that omits what the author needed, in which case the loop is the symptom
again.
> camel-jbang-mcp - a repeated identical tool call should answer differently,
> so an agent can notice it is looping
> ----------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25075
> URL: https://issues.apache.org/jira/browse/CAMEL-25075
> Project: Camel
> Issue Type: Improvement
> Components: camel-jbang
> Reporter: Claus Ibsen
> Priority: Major
>
> An agent that asks the same tool the same question over and over gets the
> same answer each time, with nothing in the answer to tell it so. A local
> model does this, and it is the largest single cause of failure on one rung of
> the benchmark.
> Measured on the {{connect-service-sql}} example, whose first step writes a
> route that inserts a row with named parameters. Across 28 attempts in three
> 10-pass runs, the calls of a failing attempt look like this (s19-2, verbatim):
> {noformat}
> camel_get_files()
> camel_get_files(sql.camel.yaml)
> camel_get_files(application.properties)
> camel_get_files(orders/order-1001.json)
> camel_catalog_doc(sql)
> camel_catalog_doc(sql)
> ... 16 times in total, identical arguments
> {noformat}
> It never wrote a file, and the step failed with its whole budget spent. Of 10
> passes of that step, 8 ended pinned at the tool-call ceiling, and in most of
> them the majority of the calls were identical repeats of
> {{camel_catalog_doc(sql)}}. Raising the ceiling does not help: the same step
> was run at 12, 20 and 24 calls, and at 24 the extra calls were spent on more
> repeats.
> What it was hunting for was the SQL dialect of the database, which the
> catalog cannot supply, so the answer was never going to change. But a model
> has no way of noticing that the answer it just received is the one it already
> had.
> This is fixable outside the model. Some options, in rough order of how little
> they assume:
> * the server notices that a call repeats an earlier call of the same session
> with the same arguments, and answers with something different: that the
> answer is unchanged, how many times it has been asked, and what else is worth
> trying ({{camel_catalog_docs}} for the page, {{camel_catalog_sample}} for a
> shape, {{camel_get_files}} for what is on disk);
> * the repeat answer carries a short list of the tools not yet used in this
> session, since an agent going in circles is usually one that has not thought
> of the next tool;
> * a cheaper variant: the answer is returned as before with one line prepended
> saying it is a repeat.
> The point is not to refuse the call -- a legitimate re-read after an edit
> must still work -- but to make a repeat *look different* so it can break the
> loop.
> Filed from the local-model benchmark of the camel-jbang-mcp server; see
> CAMEL-25040 for the related finding that the catalog answer omitted what the
> author needed in the first place.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)