jnioche commented on PR #2213: URL: https://github.com/apache/stormcrawler/pull/2213#issuecomment-5916809566
might be useful for #2212 obviously, this does not test how we get the URLs from the backend but given the write-heavy nature of crawling, the status updater bolts are often the bottleneck. I will share some stats later. `docker run -d --rm --name opensearch-bench -p 127.0.0.1:9200:9200 -e discovery.type=single-node -e DISABLE_SECURITY_PLUGIN=true -e bootstrap.memory_lock=true -e "OPENSEARCH_JAVA_OPTS=-Xms4g -Xmx4g" --ulimit memlock=-1:-1 --ulimit nofile=65536:65536 opensearchproject/opensearch:3.7.0` make sure we have a well configured opensearch-conf file then `~/apache-storm-3.0.0/bin/storm local target/ethicrawl-0.1-SNAPSHOT.jar org.apache.stormcrawler.persistence.StatusUpdaterBenchmark org.apache.stormcrawler.persistence.MemoryStatusUpdater opensearch-conf.yaml crawler-conf.yaml /home/julien/url-frontier/outlinks.ndjson` It will be interesting to see if the `opensearch-java` module is faster than `opensearch` and how either of them compare against URLFrontier. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
