[ 
https://issues.apache.org/jira/browse/SOLR-18307?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121783#comment-18121783
 ] 

David Smiley commented on SOLR-18307:
-------------------------------------

Excellent Mikhail!  I used this today on a fairly realistic data set at work, 
and found it indeed outperforms the built-in join handily.  I have something I 
shared in the last meetup that I have not yet publicly shared.  Based on the 
numbers I'm seeing across the 3, this is what I found (again, my query, my data 
– YMMV):

Measured on one shard replica (~1.03M docs, 19 segments), `distrib=false`, 
`cache=false`, query: _redacted_
|| ||{!compositeNested} + \{!child} + \{!parent}x3||{!join method=topLevelDV} 
×3||{!auxIndexJoin} ×3 ||
||Warm query time|*~6 ms*|~190 ms|~12-15 ms|
||One-time cost|~430 ms to build the merged view (250-500 ms range, 4 
builds)|~500 ms first query (top-level ordinal map)|1.5-2 s first query|
||What the one-time cost covers|All queries until the next commit (independent 
of the query)|All joins on that field until the next commit|Only the segments 
the inner query reaches; a new query touching other segments paid 1.5 s again|
||Can be moved to commit-time warming|Yes, one warming query|Yes|Yes, but the 
warming query must match every doc on the inner side for full coverage (not yet 
measured)|
||Index or schema changes|Use nested docs + index sorting.  Needs a 
concatenated  nest path field.|None|Sidecar index on disk beside the main index|
||Status|Prototype|Stock Solr|Experimental in Solr 10|

Notes:
 - All three return the same result. The join variants link entries to their 
parent by `_parent_document_id`. The ` \{!compositeNested}` variant uses block 
joins over a merged view where each doc's entries are nested under it.

 - In this run, nothing was warmed at commit, so every one-time cost was paid 
by the first user query after a commit.
 - ` \{!auxIndexJoin}` ran with Solr's search executor threads at their default 
(-1). Its full-coverage warming cost is still to be measured.

🤖 This comment was *partially* generated by AI (Claude Code).

> segment-parallel query time join with sidecar index 
> ----------------------------------------------------
>
>                 Key: SOLR-18307
>                 URL: https://issues.apache.org/jira/browse/SOLR-18307
>             Project: Solr
>          Issue Type: Improvement
>          Components: query
>            Reporter: Mikhail Khludnev
>            Assignee: Mikhail Khludnev
>            Priority: Major
>              Labels: pull-request-available
>             Fix For: 10.1
>
>         Attachments: Screenshot from 2026-08-29 18-14-22.png, hotspot.md
>
>          Time Spent: 50m
>  Remaining Estimate: 0h
>
> Hereby I propose a new query-time join implementation
> * it's segment parallel 
> * it's lazy - utilizes Two-Phase searching and beneficial from highly 
> selective filter in "to"-side, as well as leap over whole "to" segments. 
> * uses Lucene DocValues in sidecar - just an addition, there's no hard 
> surgery on existing codebase. It's just a QParserPlugin.  
> * joins crosscore, and works in cloud mode
> * may use Memory Directory for sidecar
> * writes sidecar lazily, but you can write it ahead
> * sweeps segments which are not needed anymore 
> * the [simple benchmark|https://github.com/mkhludnev/aijoin-benchmark] 
> demonstrates a few times gain 
> WIP https://github.com/apache/solr/pull/4749
> benchmark: https://github.com/mkhludnev/aijoin-benchmark
> Caveat: during the demo 7/15/26 we evidence nearly the same size of sidecar 
> index. it should be further investigated. 
> Note: the name comes from Auxiliary Index Join. 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to