Andrea Cosentino created CAMEL-24784:
----------------------------------------

             Summary: camel-opa - the REST decision call has no timeout and 
builds a new HttpClient per message
                 Key: CAMEL-24784
                 URL: https://issues.apache.org/jira/browse/CAMEL-24784
             Project: Camel
          Issue Type: Bug
            Reporter: Andrea Cosentino
            Assignee: Andrea Cosentino


h2. Problem

In {{evaluationMode=rest}} - the default - the decision call cannot time out, 
and allocates a fresh HTTP client for every message.

{{camel-opa}} builds its client with {{new OPAClient(serverUrl)}} / {{new 
OPAClient(serverUrl, headers)}}. Neither passes an {{HTTPClient}}, so the SDK 
falls back to {{SpeakeasyHTTPClient}}, whose entire implementation is:

{code:java}
public HttpResponse<InputStream> send(HttpRequest request) throws ... {
    HttpClient client = HttpClient.newHttpClient();
    ...
    return client.send(request, HttpResponse.BodyHandlers.ofInputStream());
}
{code}

Verified against the {{com.styra:opa:2.1.1}} bytecode.

h3. 1. No timeout

{{HttpClient.newHttpClient()}} applies no connect timeout, and nothing in the 
SDK sets {{HttpRequest.Builder.timeout}} (checked across every class in the 
artifact). Both JDK defaults are "wait indefinitely".

So an OPA server that accepts the connection and then does not answer parks the 
calling thread for ever. For a component whose contract is *fail closed*, this 
is the worst available failure mode: it never reaches a decision, so it never 
denies - it simply stops. {{failOpen}} does not help either, because that 
branch is downstream of the call that never returns.

A *refused* connection fails fast, so this only appears against a host that 
accepts and stalls - precisely the case that does not reproduce locally.

{{evaluationMode=wasm}} does not have this problem: CAMEL-24741 gave it a 
bounded {{borrowTimeout}}. The gap is on the default mode.

h3. 2. A new HttpClient per message

{{send()}} constructs an {{HttpClient}} per request. On the Java 17 baseline 
{{HttpClient}} is not {{AutoCloseable}}, so each one holds its selector thread 
and executor until it is collected. {{OpaProducerHealthCheck}} already shares a 
single client for exactly this reason, and the probe fires once per health poll 
- while this fires once per *message* through an {{OpaSecurityPolicy}}.

h2. Fix

{{OPAClient}} has a constructor {{camel-opa}} never uses: {{OPAClient(String 
serverUrl, HTTPClient)}}. A small camel-owned {{HTTPClient}} holding one shared 
{{java.net.http.HttpClient}} (built with a connect timeout) and rebuilding each 
request with a read timeout closes both. {{HttpRequest.newBuilder(HttpRequest, 
BiPredicate)}} makes the rebuild possible without reconstructing the request by 
hand - confirmed on 17.

New options {{connectionTimeout}} (default 10s) and {{requestTimeout}} (default 
30s). A borrow that times out is an evaluation failure, not a deny, so it fails 
closed - or proceeds under {{failOpen}} - like any other.

h2. Scope

{{main}} only. {{camel-opa}} is new and unreleased in 4.23.0, so nothing 
released changes and no advisory is warranted.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to