[ 
https://issues.apache.org/jira/browse/CAMEL-24243?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Work on CAMEL-24243 started by Andrea Cosentino.
------------------------------------------------
> camel-aws2-athena: a still-running query is relaunched when waitTimeout 
> expires and maxAttempts > 1
> ---------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-24243
>                 URL: https://issues.apache.org/jira/browse/CAMEL-24243
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-aws
>            Reporter: Andrea Cosentino
>            Assignee: Andrea Cosentino
>            Priority: Major
>             Fix For: 4.14.9, 4.18.4, 4.22.0
>
>
> h3. Problem
> {{Athena2Producer.startQueryExecution()}} drives two nested loops:
> {code:java}
> while (athena2QueryHelper.shouldAttempt()) {
>     queryExecutionId = doStartQueryExecution(athenaClient, 
> exchange).queryExecutionId();
>     athena2QueryHelper.markAttempt();
>     while (athena2QueryHelper.shouldWait()) {
>         athena2QueryHelper.doWait();
>         getQueryExecutionResponse = doGetQueryExecution(queryExecutionId, 
> athenaClient);
>         athena2QueryHelper.setStatusFrom(getQueryExecutionResponse);
>     }
> }
> {code}
> The inner loop exits on *three* different conditions: success, failure/retry, 
> and {{waitTimeout}} expiry. Only the first two set state on the helper -- 
> {{setStatusFrom()}} assigns {{isSuccess}}/{{isFailure}}/{{isRetry}} only when 
> {{isComplete()}} is true.
> When the inner loop exits because {{millisWaited >= waitTimeout}} while the 
> query is still {{QUEUED}} or {{RUNNING}}, none of those flags are set. 
> {{shouldAttempt()}} then sees {{attempts < maxAttempts}}, no failure, no 
> success, not interrupted -- and returns {{true}}. The outer loop calls 
> {{doStartQueryExecution}} again, *submitting a brand-new Athena query* while 
> the previous one is still running.
> h3. Impact
> With {{maxAttempts > 1}} (the documented way to retry failed queries):
> * every {{waitTimeout}} expiry submits another execution of the same SQL;
> * the earlier executions are orphaned -- still running, still scanning data, 
> still billed;
> * the {{CamelAwsAthenaQueryExecutionId}} header returned to the route belongs 
> to the *last* submission, so the caller cannot correlate or cancel the 
> earlier ones;
> * {{clientRequestToken}} is not auto-generated 
> ({{Athena2Producer.determineClientRequestToken}} reads only a header or the 
> endpoint option), so Athena does not deduplicate the submissions by default.
> This contradicts the documented contract in {{aws2-athena-component.adoc}}:
> bq. Upon failure, the query would be automatically retried up to two more 
> times if the failure state indicates the query may succeed upon retry
> Retry is documented as a response to *failure states*, not to a query simply 
> taking longer than {{waitTimeout}}.
> h3. Proposed fix
> Treat {{waitTimeout}} expiry as a terminal condition for the attempt loop: a 
> query that has not completed within the wait window should be reported back 
> as-is (its last known state and {{queryExecutionId}}) rather than being 
> resubmitted. Only a completed-and-retryable failure should consume another 
> attempt.
> h3. Affected versions
> The loop is identical on {{main}}, {{camel-4.18.x}} and {{camel-4.14.x}}.
> ----
> _Reported by Claude Code on behalf of acosentino, from an automated audit of 
> the camel-aws components._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to