jnioche commented on PR #2213:
URL: https://github.com/apache/stormcrawler/pull/2213#issuecomment-5916809566

   might be useful for #2212 
   
   obviously, this does not test how we get the URLs from the backend but given 
the write-heavy nature of crawling, the status updater bolts are often the 
bottleneck.
   
   I will share some stats later. 
   
   `docker run -d --rm --name opensearch-bench -p 127.0.0.1:9200:9200 -e 
discovery.type=single-node -e DISABLE_SECURITY_PLUGIN=true -e 
bootstrap.memory_lock=true -e "OPENSEARCH_JAVA_OPTS=-Xms4g -Xmx4g" --ulimit 
memlock=-1:-1 --ulimit nofile=65536:65536 opensearchproject/opensearch:3.7.0`
   
   make sure we have a well configured opensearch-conf file then 
   
   `~/apache-storm-3.0.0/bin/storm  local target/ethicrawl-0.1-SNAPSHOT.jar 
org.apache.stormcrawler.persistence.StatusUpdaterBenchmark  
org.apache.stormcrawler.persistence.MemoryStatusUpdater opensearch-conf.yaml 
crawler-conf.yaml /home/julien/url-frontier/outlinks.ndjson`
   
   It will be interesting to see if the `opensearch-java` module is faster than 
`opensearch` and how either of them compare against URLFrontier.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to