jnioche opened a new pull request, #2225:
URL: https://github.com/apache/stormcrawler/pull/2225

   The `StatusUpdaterBolt` of the `opensearch-java` module is slower than the 
one of the legacy `opensearch` module. A likely cause is the way 
`AsyncBulkProcessor` executes the bulk requests.
   
   It ran the blocking `OpenSearchClient.bulk()` call on a `ThreadPoolExecutor` 
bounded to `concurrentRequests` threads, with a `SynchronousQueue` and 
`CallerRunsPolicy`. The semaphore permit was released *before* the listener 
processed the response, so the bolt could acquire it and submit the next 
request while the only worker (with the default `concurrentRequests=1`) was 
still busy in `afterBulk`. The task was then rejected and the bulk call ran on 
the bolt's own executor thread, blocking it for the whole HTTP round trip 
instead of letting it fill the next batch.
   
   ### Changes
   - `AsyncBulkProcessor` uses `OpenSearchAsyncClient.bulk()`, like the legacy 
`BulkProcessor` used `bulkAsync`; the dedicated executor is removed.
   - The permit is released only once the listener has run, as in the legacy 
implementation. Back-pressure is unchanged: `add()` blocks only when 
`concurrentRequests` bulks are in flight.
   - `OpenSearchConnection` builds the async client on the same transport as 
the sync one.
   - Tests updated, plus a new test checking that `add()` does not wait for the 
bulk response.
   
   Note that `afterBulk` now runs on the HTTP client's I/O reactor thread, 
which is also where the legacy module ran it.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to