This is an automated email from the ASF dual-hosted git repository.
yiguolei pushed a commit to branch branch-4.1
in repository https://gitbox.apache.org/repos/asf/doris.git
The following commit(s) were added to refs/heads/branch-4.1 by this push:
new 089fa51f229 [fix](lance) Select metric-compatible vector index (#67553)
089fa51f229 is described below
commit 089fa51f22942fdedcb78bed999818742de020b2
Author: FANNG <[email protected]>
AuthorDate: Wed Sep 16 10:34:09 2026 +0800
[fix](lance) Select metric-compatible vector index (#67553)
### What problem does this PR solve?
Issue Number: close #67117
Related PR: None
Problem Summary: When multiple Lance vector indexes exist on the same
column, Doris previously selected the first physical index segment
before checking its metric. If that logical index used a different
metric, Doris fell back to flat search without trying a later compatible
index.
Group physical index segments by logical index name, visit logical
indexes in lexicographic name order, filter each complete group by the
requested metric, and select the first group that can be planned safely.
If no compatible usable group exists, preserve the existing flat-search
fallback. This keeps every physical segment of one logical index
together and never mixes segment UUIDs from different logical indexes.
### Release note
Fix Lance vector index selection when multiple indexes with different
metrics exist on the same column.
### Check List (For Author)
- Test: Unit Test / FE validation
- `mvn validate -pl fe-core -am -DskipTests` passed, including
Checkstyle.
- Added `LanceScanNodeTest` cases for metric matching, metadata-order
independence, lexicographic logical-index selection, DEFAULT-to-L2
behavior, unsafe-group skipping, and multi-segment UUID grouping.
- `./run-fe-ut.sh --run
org.apache.doris.datasource.lance.source.LanceScanNodeTest` was
attempted but blocked before `fe-core`: the macOS prebuilt thirdparty
Thrift compiler generates Java incompatible with this historical branch,
causing `fe-common` compilation errors.
- Behavior changed: Yes. Doris now tries metric-compatible logical Lance
indexes in lexicographic index-name order before falling back to flat
search.
- Does this need documentation: No
---
.../scripts/lance_build_preinstalled_catalog.py | 18 ++-
.../datasource/lance/source/LanceScanNode.java | 48 ++++----
.../datasource/lance/source/LanceScanNodeTest.java | 128 +++++++++++++++++++++
.../test_lance_vector_search_index_matrix.groovy | 9 +-
4 files changed, 161 insertions(+), 42 deletions(-)
diff --git
a/docker/thirdparties/docker-compose/iceberg/scripts/lance_build_preinstalled_catalog.py
b/docker/thirdparties/docker-compose/iceberg/scripts/lance_build_preinstalled_catalog.py
index 92220554c7b..deca56d4f6e 100644
---
a/docker/thirdparties/docker-compose/iceberg/scripts/lance_build_preinstalled_catalog.py
+++
b/docker/thirdparties/docker-compose/iceberg/scripts/lance_build_preinstalled_catalog.py
@@ -388,7 +388,6 @@ VECTOR_TABLES = {
},
}
-
# ---------------------------------------------------------------------------
# Breadth tier
# ---------------------------------------------------------------------------
@@ -397,16 +396,13 @@ VECTOR_TABLES = {
# reached the index. That costs roughly 190KB per cell, so the tier
deliberately covers one
# representative cell per axis rather than the whole matrix.
#
-# The breadth tier covers everything the depth tier leaves out, at plan level
only. It is a
-# single table carrying one vector column per remaining cell, each with
exactly one index -
-# one column per cell, never several indexes on one column, because only the
first index
-# built on a column is reachable. Measured on Lance: with a cosine and a dot
index on one
-# column, whichever was created first answers its metric from the index and
the other falls
-# back to a silent brute-force scan. Doris lands in the same place by a
different route -
-# LanceScanNode.selectIndexSegments keeps only the segments of the first index
it finds for
-# the column's field id, so the second index is invisible to the planner and
metricMatches
-# then rejects the query whose metric it does not carry. Either way a column
is the unit that
-# can hold a testable index, and 64 rows is enough to train one.
+# The breadth tier covers every remaining cell at plan level only.
+# It is a single table carrying one vector column per remaining cell, each
with exactly one
+# index. Keeping one index per column isolates every type x metric x algorithm
cell;
+# same-column multi-index selection is covered by LanceScanNodeTest's metadata
fixtures. A
+# column is not limited to one logical vector index: Doris groups physical
segments by index
+# name and picks the first lexicographic group that matches the requested
metric and can safely
+# plan splits. Sixty-four rows are enough to train each matrix index.
#
# What this tier proves is narrower than the depth tier's, and the
documentation must not
# conflate them: it shows Doris plans an indexed split and the backend answers
it, NOT that
diff --git
a/fe/fe-core/src/main/java/org/apache/doris/datasource/lance/source/LanceScanNode.java
b/fe/fe-core/src/main/java/org/apache/doris/datasource/lance/source/LanceScanNode.java
index 5b079ab4e85..7c0b90f26b1 100644
---
a/fe/fe-core/src/main/java/org/apache/doris/datasource/lance/source/LanceScanNode.java
+++
b/fe/fe-core/src/main/java/org/apache/doris/datasource/lance/source/LanceScanNode.java
@@ -65,6 +65,7 @@ import java.util.Locale;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
+import java.util.TreeMap;
import java.util.UUID;
/**
@@ -384,23 +385,25 @@ public class LanceScanNode extends FileQueryScanNode {
+ "' has no field ID in the Lance schema");
}
- List<LanceIndexSegmentInfo> matchingSegments =
selectVectorIndexSegments(
- metadata.getIndexSegments(), searchFieldId);
- if (matchingSegments.isEmpty() || !metricMatches(vectorSearchParam,
matchingSegments)) {
- return Optional.empty();
- }
+ for (List<LanceIndexSegmentInfo> indexSegments :
selectVectorIndexSegmentGroups(
+ metadata.getIndexSegments(), searchFieldId)) {
+ if (!metricMatches(vectorSearchParam, indexSegments)) {
+ continue;
+ }
- Optional<IndexSegmentSplitPlan> indexPlan = planIndexSegments(
- metadata, matchingSegments, visibleFragments, false);
- if (!indexPlan.isPresent()) {
- return Optional.empty();
+ Optional<IndexSegmentSplitPlan> indexPlan = planIndexSegments(
+ metadata, indexSegments, visibleFragments, false);
+ if (!indexPlan.isPresent()) {
+ continue;
+ }
+ IndexSegmentSplitPlan plan = indexPlan.get();
+ plannedIndexSegments = plan.splitCount();
+ plannedIndexFragments = plan.indexSegmentFragmentCount();
+ plannedUnindexedFragments = plannedFragments -
plannedIndexFragments;
+ appendUnindexedFragmentSplits(plan, visibleFragments);
+ return Optional.of(plan.buildSplits());
}
- IndexSegmentSplitPlan plan = indexPlan.get();
- plannedIndexSegments = plan.splitCount();
- plannedIndexFragments = plan.indexSegmentFragmentCount();
- plannedUnindexedFragments = plannedFragments - plannedIndexFragments;
- appendUnindexedFragmentSplits(plan, visibleFragments);
- return Optional.of(plan.buildSplits());
+ return Optional.empty();
}
private List<Split> createFullTextIndexSegmentSplits(LanceTableMetadata
metadata,
@@ -436,22 +439,17 @@ public class LanceScanNode extends FileQueryScanNode {
return plan.buildSplits();
}
- private static List<LanceIndexSegmentInfo> selectVectorIndexSegments(
+ private static List<List<LanceIndexSegmentInfo>>
selectVectorIndexSegmentGroups(
List<LanceIndexSegmentInfo> indexSegments, int fieldId) {
- List<LanceIndexSegmentInfo> selectedSegments = new ArrayList<>();
- String selectedIndexName = null;
+ // A stable order keeps index selection independent of Lance metadata
ordering.
+ Map<String, List<LanceIndexSegmentInfo>> groupsByName = new
TreeMap<>();
for (LanceIndexSegmentInfo segment : indexSegments) {
if (!segment.isVectorIndex() ||
!segment.getFieldIds().contains(fieldId)) {
continue;
}
- if (selectedIndexName == null) {
- selectedIndexName = segment.getIndexName();
- }
- if (selectedIndexName.equals(segment.getIndexName())) {
- selectedSegments.add(segment);
- }
+ groupsByName.computeIfAbsent(segment.getIndexName(), ignored ->
new ArrayList<>()).add(segment);
}
- return selectedSegments;
+ return new ArrayList<>(groupsByName.values());
}
private static List<LanceIndexSegmentInfo> selectFullTextIndexSegments(
diff --git
a/fe/fe-core/src/test/java/org/apache/doris/datasource/lance/source/LanceScanNodeTest.java
b/fe/fe-core/src/test/java/org/apache/doris/datasource/lance/source/LanceScanNodeTest.java
index bdcf522c54a..99636231145 100644
---
a/fe/fe-core/src/test/java/org/apache/doris/datasource/lance/source/LanceScanNodeTest.java
+++
b/fe/fe-core/src/test/java/org/apache/doris/datasource/lance/source/LanceScanNodeTest.java
@@ -442,6 +442,134 @@ public class LanceScanNodeTest {
assertSplit(splits.get(1), 2, 8, 88);
}
+ @Test
+ public void testExternalSearchSelectsLaterMetricCompatibleIndex() throws
Exception {
+ UUID l2Segment =
UUID.fromString("11111111-1111-1111-1111-111111111111");
+ UUID cosineSegment =
UUID.fromString("22222222-2222-2222-2222-222222222222");
+ LanceTableMetadata metadata = LanceTableMetadata.withIndexSegments(
+ "s3://bucket/table.lance",
+ 42,
+ vectorSchema(),
+ Arrays.asList(new LanceFragmentInfo(1, 8, 8), new
LanceFragmentInfo(2, 7, 7)),
+ Collections.singletonMap("vector", 9),
+ Arrays.asList(
+ new LanceIndexSegmentInfo(l2Segment, "a_l2",
Collections.singletonList(9),
+ Arrays.asList(1L, 2L), IndexType.VECTOR, "L2"),
+ new LanceIndexSegmentInfo(cosineSegment, "z_cosine",
Collections.singletonList(9),
+ Arrays.asList(1L, 2L), IndexType.VECTOR,
"COSINE")),
+ Collections.emptyMap());
+ TExternalSearchRequest request = vectorSearchRequest(5, 0);
+
request.getSearchQuery().getVectorSearch().setMetric(TVectorMetric.COSINE);
+
+ List<Split> splits = newSearchNode(metadata, request).getSplits(2);
+
+ Assert.assertEquals(1, splits.size());
+ assertIndexSplit(splits.get(0), cosineSegment, Arrays.asList(1L, 2L),
15, 100);
+ }
+
+ @Test
+ public void
testExternalSearchSelectsMetricCompatibleIndexRegardlessOfMetadataOrder()
throws Exception {
+ UUID cosineSegment =
UUID.fromString("33333333-3333-3333-3333-333333333333");
+ UUID l2Segment =
UUID.fromString("44444444-4444-4444-4444-444444444444");
+ LanceTableMetadata metadata = LanceTableMetadata.withIndexSegments(
+ "s3://bucket/table.lance",
+ 42,
+ vectorSchema(),
+ Arrays.asList(new LanceFragmentInfo(1, 8, 8), new
LanceFragmentInfo(2, 7, 7)),
+ Collections.singletonMap("vector", 9),
+ Arrays.asList(
+ new LanceIndexSegmentInfo(cosineSegment, "z_cosine",
Collections.singletonList(9),
+ Arrays.asList(1L, 2L), IndexType.VECTOR,
"COSINE"),
+ new LanceIndexSegmentInfo(l2Segment, "a_l2",
Collections.singletonList(9),
+ Arrays.asList(1L, 2L), IndexType.VECTOR,
"L2")),
+ Collections.emptyMap());
+ TExternalSearchRequest request = vectorSearchRequest(5, 0);
+
request.getSearchQuery().getVectorSearch().setMetric(TVectorMetric.COSINE);
+
+ List<Split> splits = newSearchNode(metadata, request).getSplits(2);
+
+ Assert.assertEquals(1, splits.size());
+ assertIndexSplit(splits.get(0), cosineSegment, Arrays.asList(1L, 2L),
15, 100);
+ }
+
+ @Test
+ public void testExternalSearchSelectsSameMetricIndexByName() throws
Exception {
+ UUID laterSegment =
UUID.fromString("55555555-5555-5555-5555-555555555555");
+ UUID firstSegment =
UUID.fromString("66666666-6666-6666-6666-666666666666");
+ UUID secondSegment =
UUID.fromString("77777777-7777-7777-7777-777777777777");
+ LanceTableMetadata metadata = LanceTableMetadata.withIndexSegments(
+ "s3://bucket/table.lance",
+ 42,
+ vectorSchema(),
+ Arrays.asList(
+ new LanceFragmentInfo(1, 8, 8),
+ new LanceFragmentInfo(2, 7, 7),
+ new LanceFragmentInfo(3, 6, 6)),
+ Collections.singletonMap("vector", 9),
+ Arrays.asList(
+ new LanceIndexSegmentInfo(laterSegment, "z_l2",
Collections.singletonList(9),
+ Arrays.asList(1L, 2L, 3L), IndexType.VECTOR,
"L2"),
+ new LanceIndexSegmentInfo(firstSegment, "a_l2",
Collections.singletonList(9),
+ Collections.singletonList(1L),
IndexType.VECTOR, "L2"),
+ new LanceIndexSegmentInfo(secondSegment, "a_l2",
Collections.singletonList(9),
+ Collections.singletonList(2L),
IndexType.VECTOR, "L2")),
+ Collections.emptyMap());
+
+ List<Split> splits = newSearchNode(metadata, vectorSearchRequest(5,
0)).getSplits(2);
+
+ Assert.assertEquals(3, splits.size());
+ assertIndexSplit(splits.get(0), firstSegment,
Collections.singletonList(1L), 8, 100);
+ assertIndexSplit(splits.get(1), secondSegment,
Collections.singletonList(2L), 8, 88);
+ assertSplit(splits.get(2), 3, 8, 75);
+ }
+
+ @Test
+ public void testExternalSearchDefaultMetricSelectsL2Index() throws
Exception {
+ UUID l2Segment =
UUID.fromString("99999999-9999-9999-9999-999999999999");
+ LanceTableMetadata metadata = LanceTableMetadata.withIndexSegments(
+ "s3://bucket/table.lance",
+ 42,
+ vectorSchema(),
+ Collections.singletonList(new LanceFragmentInfo(1, 8, 8)),
+ Collections.singletonMap("vector", 9),
+ Collections.singletonList(
+ new LanceIndexSegmentInfo(l2Segment, "l2",
Collections.singletonList(9),
+ Collections.singletonList(1L),
IndexType.VECTOR, "L2")),
+ Collections.emptyMap());
+ TExternalSearchRequest request = vectorSearchRequest(5, 0);
+
request.getSearchQuery().getVectorSearch().setMetric(TVectorMetric.DEFAULT);
+
+ List<Split> splits = newSearchNode(metadata, request).getSplits(1);
+
+ Assert.assertEquals(1, splits.size());
+ assertIndexSplit(splits.get(0), l2Segment,
Collections.singletonList(1L), 8, 100);
+ }
+
+ @Test
+ public void testExternalSearchSkipsIndexGroupWithoutFragmentBitmap()
throws Exception {
+ UUID incompleteSegment =
UUID.fromString("77777777-7777-7777-7777-777777777777");
+ UUID cosineSegment =
UUID.fromString("88888888-8888-8888-8888-888888888888");
+ LanceTableMetadata metadata = LanceTableMetadata.withIndexSegments(
+ "s3://bucket/table.lance",
+ 42,
+ vectorSchema(),
+ Arrays.asList(new LanceFragmentInfo(1, 8, 8), new
LanceFragmentInfo(2, 7, 7)),
+ Collections.singletonMap("vector", 9),
+ Arrays.asList(
+ new LanceIndexSegmentInfo(incompleteSegment,
"a_incomplete",
+ Collections.singletonList(9), null,
IndexType.VECTOR, "COSINE"),
+ new LanceIndexSegmentInfo(cosineSegment, "z_cosine",
Collections.singletonList(9),
+ Arrays.asList(1L, 2L), IndexType.VECTOR,
"COSINE")),
+ Collections.emptyMap());
+ TExternalSearchRequest request = vectorSearchRequest(5, 0);
+
request.getSearchQuery().getVectorSearch().setMetric(TVectorMetric.COSINE);
+
+ List<Split> splits = newSearchNode(metadata, request).getSplits(2);
+
+ Assert.assertEquals(1, splits.size());
+ assertIndexSplit(splits.get(0), cosineSegment, Arrays.asList(1L, 2L),
15, 100);
+ }
+
@Test
public void
testExternalSearchRejectsMissingFieldIdForIndexSegmentPlanning() {
LanceTableMetadata metadata = LanceTableMetadata.withIndexSegments(
diff --git
a/regression-test/suites/external_table_p0/lance/test_lance_vector_search_index_matrix.groovy
b/regression-test/suites/external_table_p0/lance/test_lance_vector_search_index_matrix.groovy
index 6f66fe37d12..1698e2a1075 100644
---
a/regression-test/suites/external_table_p0/lance/test_lance_vector_search_index_matrix.groovy
+++
b/regression-test/suites/external_table_p0/lance/test_lance_vector_search_index_matrix.groovy
@@ -26,12 +26,9 @@ suite("test_lance_vector_search_index_matrix",
"p0,external") {
* returns the *right* rows, and it costs roughly 190KB of committed
binary per cell.
*
* This suite is the breadth tier. doris.vs_index_matrix is a single
64-row table with one
- * vector column per remaining cell, each carrying exactly one index. One
column per cell,
- * never several indexes on one column, because only the first index built
on a column is
- * reachable: Lance answers the other one with a silent brute-force scan,
and Doris gets
- * there differently but ends up the same, since
LanceScanNode.selectIndexSegments keeps
- * only the segments of the first index it finds for the column's field id
and
- * metricMatches then rejects the query whose metric that index does not
carry.
+ * vector column per remaining cell, each carrying exactly one index. One
column per cell
+ * keeps each type x metric x algorithm assertion independent; same-column
multi-index
+ * selection is covered by LanceScanNodeTest's metadata fixtures.
*
* What this proves: Doris plans an indexed split for the cell, and a
refined indexed
* search returns exactly the rows an exhaustive scan returns. That is the
same equality
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]