[
https://issues.apache.org/jira/browse/SOLR-18484?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Mikhail Khludnev updated SOLR-18484:
------------------------------------
Description:
When {{\{!auxIndexJoin}}} got no sibling filter (including {{fq}}) there's no
sense to bother with lazy two phase iteration.
Let's just defer the decision about that up to {{AIJScorerSupplier.get()}} and
when it got {{SS.get(cost==MAX)}} means there's no sibling filter, skip
approximation just eagerly drain all columns onto to-size bitsets.
was:
Hereby I propose a new query-time join implementation
* it's segment parallel
* it's lazy - utilizes Two-Phase searching and beneficial from highly selective
filter in "to"-side, as well as leap over whole "to" segments.
* uses Lucene DocValues in sidecar - just an addition, there's no hard surgery
on existing codebase. It's just a QParserPlugin.
* joins crosscore, and works in cloud mode
* may use Memory Directory for sidecar
* writes sidecar lazily, but you can write it ahead
* sweeps segments which are not needed anymore
* the [simple benchmark|https://github.com/mkhludnev/aijoin-benchmark]
demonstrates a few times gain
WIP https://github.com/apache/solr/pull/4749
benchmark: https://github.com/mkhludnev/aijoin-benchmark
Caveat: during the demo 7/15/26 we evidence nearly the same size of sidecar
index. it should be further investigated.
Note: the name comes from Auxiliary Index Join.
> auxIndexJoin perf opt - eager drain when there's no filter
> -----------------------------------------------------------
>
> Key: SOLR-18484
> URL: https://issues.apache.org/jira/browse/SOLR-18484
> Project: Solr
> Issue Type: Improvement
> Components: query
> Reporter: Mikhail Khludnev
> Assignee: Mikhail Khludnev
> Priority: Major
> Labels: pull-request-available
> Fix For: 10.1
>
>
> When {{\{!auxIndexJoin}}} got no sibling filter (including {{fq}}) there's no
> sense to bother with lazy two phase iteration.
> Let's just defer the decision about that up to {{AIJScorerSupplier.get()}}
> and when it got {{SS.get(cost==MAX)}} means there's no sibling filter, skip
> approximation just eagerly drain all columns onto to-size bitsets.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]