[
https://issues.apache.org/jira/browse/SOLR-18497?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SOLR-18497:
----------------------------------
Labels: pull-request-available (was: )
> Concurrent RELOAD of a core that failed to load creates the core once per
> request
> ---------------------------------------------------------------------------------
>
> Key: SOLR-18497
> URL: https://issues.apache.org/jira/browse/SOLR-18497
> Project: Solr
> Issue Type: Bug
> Reporter: Serhiy Bzhezytskyy
> Priority: Major
> Labels: pull-request-available
> Time Spent: 10m
> Remaining Estimate: 0h
>
> A core that failed to load stays in {{coreInitFailures}}, and
> {{CoreContainer.reload}} then takes the branch that creates it from the
> stored descriptor. That branch waits for the per-core reservation and calls
> {{createFromDescriptor}} without checking again whether the core was loaded
> in the meantime. When several RELOAD requests arrive together after the cause
> of the failure was fixed, every one of them creates the core: the first
> succeeds, the others fail on the index lock with HTTP 500, and each failure
> is recorded as a new init failure. The core then serves requests while
> CoreAdmin STATUS still lists it under {{initFailures}}.
> To reproduce (user-managed, {{<home>/configsets/_default}} copied from the
> distribution):
> {code:bash}
> bin/solr start --user-managed --solr-home <home>
> curl '.../solr/admin/cores?action=CREATE&name=f1&configSet=_default'
> bin/solr stop
> mv <home>/configsets/_default <home>/configsets/_hidden # f1 now fails
> to load at startup
> bin/solr start --user-managed --solr-home <home>
> mv <home>/configsets/_hidden <home>/configsets/_default # the cause is
> fixed
> for i in 1 2 3 4 5 6 7 8; do curl
> '.../solr/admin/cores?action=RELOAD&core=f1' & done; wait
> {code}
> One request returns 200 and seven return 500, "Unable to create core [f1]",
> caused by
> {noformat}
> LockObtainFailedException: Index dir '.../f1/data/index/' of core 'f1' is
> already locked.
> {noformat}
> The core answers queries afterwards, but {{STATUS}} still shows f1 under
> {{initFailures}}. A single RELOAD creates the core once and leaves
> {{initFailures}} empty.
> The reload should not create the core again when another request loaded it
> while this one waited for the reservation.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]