jnioche opened a new pull request, #2225: URL: https://github.com/apache/stormcrawler/pull/2225
The `StatusUpdaterBolt` of the `opensearch-java` module is slower than the one of the legacy `opensearch` module. A likely cause is the way `AsyncBulkProcessor` executes the bulk requests. It ran the blocking `OpenSearchClient.bulk()` call on a `ThreadPoolExecutor` bounded to `concurrentRequests` threads, with a `SynchronousQueue` and `CallerRunsPolicy`. The semaphore permit was released *before* the listener processed the response, so the bolt could acquire it and submit the next request while the only worker (with the default `concurrentRequests=1`) was still busy in `afterBulk`. The task was then rejected and the bulk call ran on the bolt's own executor thread, blocking it for the whole HTTP round trip instead of letting it fill the next batch. ### Changes - `AsyncBulkProcessor` uses `OpenSearchAsyncClient.bulk()`, like the legacy `BulkProcessor` used `bulkAsync`; the dedicated executor is removed. - The permit is released only once the listener has run, as in the legacy implementation. Back-pressure is unchanged: `add()` blocks only when `concurrentRequests` bulks are in flight. - `OpenSearchConnection` builds the async client on the same transport as the sync one. - Tests updated, plus a new test checking that `add()` does not wait for the bulk response. Note that `afterBulk` now runs on the HTTP client's I/O reactor thread, which is also where the legacy module ran it. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
