Hi Bartosz, Please create a Jira ticket / PR, this looks like a solid investigation and you already did a lot of background work for a PR ;)
Andrzej > On 3 Jul 2026, at 11:38, Bartosz Fidrysiak <[email protected]> wrote: > > I prepared a benchmark (that can be applied to Solr 9 & 10 code bases). It > was executed against Solr 9.10.1 and Solr 9.10.1-SNAPSHOT (including both > enhancements). It starts a local mini Solr cluster with caching disabled and > measures the average execution time of the following requests: > qSimple (no collapse) > qCollapseWithoutSort > qCollapseByStr > qCollapseByDate > qCollapseByLong > qCollapseByDateAndStr > > In the prepared benchmark, each request is executed multiple times across the > scenarios described below. Note that documents from the same groups are > distributed evenly across all segments in this benchmark, so to trigger > cross-segment comparisons, numSegments should be set to a value greater than > one. > > numDocs, numGroups, numSegments > 2_000_000, 100_000, 1 > 2_000_000, 100_000, 10 > > > Solr 9.10.1 benchmark results > Benchmark (numDocs) (numGroups) (numSegments) > Mode Cnt Score Error Units > CollapsingSearch.collapseByDate 2000000 100000 1 > avgt 3 40.520 ? 15.037 ms/op > CollapsingSearch.collapseByDate 2000000 100000 10 > avgt 3 40.355 ? 28.083 ms/op > CollapsingSearch.collapseByDateAndStr 2000000 100000 1 > avgt 3 405.478 ? 46.023 ms/op > CollapsingSearch.collapseByDateAndStr 2000000 100000 10 > avgt 3 398.666 ? 81.423 ms/op > CollapsingSearch.collapseByLong 2000000 100000 1 > avgt 3 41.811 ? 5.349 ms/op > CollapsingSearch.collapseByLong 2000000 100000 10 > avgt 3 40.800 ? 8.423 ms/op > CollapsingSearch.collapseByStr 2000000 100000 1 > avgt 3 1862.934 ? 148.822 ms/op > CollapsingSearch.collapseByStr 2000000 100000 10 > avgt 3 1741.207 ? 34.472 ms/op > CollapsingSearch.collapseWithoutSort 2000000 100000 1 > avgt 3 9.780 ? 1.670 ms/op > CollapsingSearch.collapseWithoutSort 2000000 100000 10 > avgt 3 11.779 ? 2.800 ms/op > CollapsingSearch.simple 2000000 100000 1 > avgt 3 2.561 ? 20.237 ms/op > CollapsingSearch.simple 2000000 100000 10 > avgt 3 1.597 ? 0.104 ms/op > > > Solr 9.10.1-SNAPSHOT (with both enhancements) benchmark results > Benchmark (numDocs) (numGroups) (numSegments) > Mode Cnt Score Error Units > CollapsingSearch.collapseByDate 2000000 100000 1 > avgt 3 37.959 ? 14.352 ms/op > CollapsingSearch.collapseByDate 2000000 100000 10 > avgt 3 38.329 ? 8.270 ms/op > CollapsingSearch.collapseByDateAndStr 2000000 100000 1 > avgt 3 41.098 ? 12.300 ms/op > CollapsingSearch.collapseByDateAndStr 2000000 100000 10 > avgt 3 43.339 ? 34.562 ms/op > CollapsingSearch.collapseByLong 2000000 100000 1 > avgt 3 43.598 ? 41.345 ms/op > CollapsingSearch.collapseByLong 2000000 100000 10 > avgt 3 39.329 ? 3.095 ms/op > CollapsingSearch.collapseByStr 2000000 100000 1 > avgt 3 37.017 ? 14.269 ms/op > CollapsingSearch.collapseByStr 2000000 100000 10 > avgt 3 1880.135 ? 64.181 ms/op > CollapsingSearch.collapseWithoutSort 2000000 100000 1 > avgt 3 9.422 ? 1.844 ms/op > CollapsingSearch.collapseWithoutSort 2000000 100000 10 > avgt 3 11.580 ? 2.093 ms/op > CollapsingSearch.simple 2000000 100000 1 > avgt 3 2.542 ? 20.557 ms/op > CollapsingSearch.simple 2000000 100000 10 > avgt 3 1.599 ? 0.029 ms/op > > > > > Both enhancements (lazy string loading while creating new collapse group & > ordinal fast path optimization for docs from the same segment) bring huge > benefits in certain cases for collapse queries with string collapse sort > fields in Solr 9. > I also sent a request asking for creating a JIRA account so I can create a > performance degradation issue and provide more details there. I can also > contribute cause I already has tested patches ready. > > Kind regards, > Bartosz Fidrysiak > > On Wed, Jul 1, 2026 at 8:37 PM Bartosz Fidrysiak <[email protected] > <mailto:[email protected]>> wrote: >> We investigated the collapse-with-string-sort performance issue in depth and >> identified two practical enhancements: >> Enhancement 1: Lazy loading of string sort values for group heads — the >> string value is only materialized (triggering LZ4 decompression) when a >> competing document actually appears for the group >> Enhancement 2: Ordinal-based comparison for same-segment documents — instead >> of materializing string values, ordinals are compared directly using a >> simple integer comparison. Ordinals are numeric, segment-local, and require >> no decompression. >> >> We benchmarked both enhancements against the official Solr 9.10.1 Docker >> image: >> • Enhancement 1 alone significantly improves collapse-sort-by-date-and-str - >> Solr 9 query times are better than Solr 8, rather than being twice as bad. >> This is because string tiebreaker values no longer need to be eagerly loaded >> from sorted doc values when the winner can already be determined by the date >> comparison. However, in this snapshot, collapse-sort-by-str (string-only >> sort) shows no improvement over Solr 8 - the ordinal fast path is not yet >> active. >> • Enhancements 1 & 2 combined significantly improve both >> collapse-sort-by-date-and-str and collapse-sort-by-str. The string-only sort >> case benefits particularly well in our dataset because a large proportion of >> documents share the same segments, making ordinal comparisons widely >> applicable. >> >> We prepared also a patch with the changes (see attachments) as a proposition >> that could be introduced to Solr 9 & Solr 10. The patch contains both >> enhancements and is quite small: >> LazyStringValue (+57 / 0) - a new class. It contains materialize() method >> responsible for loading string values from sorted doc values only when it is >> really needed. >> CollapsingQParserPlugin.SortFieldsCompare (+55 / -3) - existing class used >> to create collapse group heads and comparing documents from the same group. >> TestCollapseQParserPlugin (+146 / 0) - tests >> >> What's the procedure of proposing the change? Can you create a JIRA for the >> issue? I can contribute and prepare PRs if needed. >> >> Kind regards, >> Bartosz >> >> >> On Thu, Jun 25, 2026 at 11:13 AM Bartosz Fidrysiak <[email protected] >> <mailto:[email protected]>> wrote: >>> We identified a 2–3x performance regression in Solr 9.10.1 compared to Solr >>> 8.11.2 for collapse >>> queries that use a string field as a collapse sort field. >>> >>> >>> Test setup >>> ---------- >>> >>> To measure the regression under real production conditions, we configured >>> both clusters to receive identical traffic simultaneously — every Solr >>> request is sent to both instances at the same time, making the comparison >>> direct and unbiased. Both clusters have the same number of nodes, >>> documents, shards, and shard ranges. The data is sharded by tenant ID, so >>> each request is served by a single shard with no cross-shard overhead. Solr >>> schema is the same for both clusters. >>> >>> We tested six query variants covering different combinations of collapse >>> sort fields: no collapse, collapse with date sort, date+long sort, >>> date+string sort, and string-only sort (see attachments). The results show >>> that queries with a string field in the collapse sort are consistently and >>> significantly slower in Solr 9, while queries using only numeric or date >>> sort fields show no regression. Notably, the string field used in the >>> collapse sort has very high cardinality, and the worst-case queries process >>> millions of documents. >>> >>> >>> >>> >>> >>> >>> Root cause >>> ---------- >>> >>> JFR profiling of the worst-case query (sort="modified_date desc, >>> document_id asc", ~7M documents) confirmed the root cause. >>> >>> >>> Lucene 9 changed the internal format for SortedDocValues >>> (Lucene90DocValuesProducer). The term dictionary (TermsDict) now stores >>> string values in LZ4-compressed blocks. In Lucene 8, the same data was held >>> uncompressed in direct memory — reads were instant. In Lucene 9, every time >>> the collapse logic needs to materialize a string value for comparison or to >>> record a new group winner, it must decompress an LZ4 block. For ~7M >>> documents, this decompression is triggered on nearly every document via the >>> following call chain: >>> >>> SortFieldsCompare >>> -> TermOrdValLeafComparator.copy() >>> -> lookupOrd() >>> -> TermsDict.decompressBlock() >>> -> LZ4.decompress() >>> >>> LZ4 decompression accounts for almost 40% of CPU time in the query-serving >>> thread in Solr 9, >>> versus near zero in Solr 8. >>> >>> Similar concerns were raised in >>> https://github.com/apache/lucene/issues/11485 >>> >>> Findings >>> --------- >>> Lucene90DocValuesProducer is used in both main Solr sort and collapse sort >>> in a different way. >>> The main sort in Solr queries uses two-phase comparison: ordinals first if >>> both values are in the same segment, and only materializes the string value >>> via lookupOrd() if they reside in different segments. This is not the case >>> fort collapse sort. >>> >>> During collapse with sort by string field, Solr compares candidates against >>> the current group winner via SortFieldsCompare, which always calls copy() >>> on the comparator for every document - regardless of whether the document >>> and the current group winner are from the same segment or different >>> segments. The same copy() call triggers lookupOrd() and LZ4 decompression, >>> and also stores the string value as the new group winner if the document >>> wins the comparison. There is no ordinal-only shortcut. >>> >>> Questions >>> --------- >>> >>> Q1: What are your recommendations for improving the performance of collapse >>> queries that use a string field as a sort tiebreaker in Solr 9? >>> >>> Q2: Is it possible to disable LZ4 compression for SortedDocValues term >>> dictionaries — either via a configuration property or a docValuesFormat >>> option — or is this something that could be planned for a future release? >>> >>> Q3: Would it be feasible to lazily materialize string field values in >>> CollapsingQParserPlugin for group winners, so that lookupOrd() is only >>> called when a cross-segment comparison is actually needed? This could >>> improve performance for queries where most groups contain only one document >>> or when two documents reside in the same segment. >>> >>> Kind regards, >>> Bartosz
