morningman opened a new issue, #66418: URL: https://github.com/apache/doris/issues/66418
This issue summarizes what Apache Doris supports today for the three open table/data formats at the center of multi-modal lakehouse workloads — **Apache Iceberg**, **Apache Paimon**, and **Lance** — and what is planned for the two releases scheduled in the second half of 2026. It also introduces two new data sources arriving in 5.0: **ADBC** and **Apache Fluss**. The goal is to give users a single place to see the capability matrix and the roadmap, so they can plan adoption against concrete versions rather than guessing. > Everything marked for 4.2 and 5.0 is in active development. Scope may still change before release. Feedback and use cases are very welcome — please comment below. **Legend:** ✅ Supported · 🔬 Supported, marked experimental · 🆕 New in 4.2 · 📅 Planned for 5.0 · — Not applicable, or no current plan ## Release Timeline | Version | Date | Status | Lakehouse highlights | |---|---|---|---| | 4.1 | Released | GA | Iceberg write enhancements, Deletion Vector reads, partial V3 support | | 4.2 | September 2026 | In development | Variant support for Iceberg and Paimon, Paimon write support, Lance read support | | 5.0 | November 2026 | In development | ADBC Catalog, Fluss Catalog, Paimon index reads | ## Overview | Data source | Supported today (≤ 4.1) | 4.2 (2026-09) | 5.0 (2026-11) | |---|---|---|---| | **Iceberg** | Full read/write: all Catalog types, V1/V2 formats, Time Travel and system tables, INSERT / OVERWRITE / CTAS, DELETE / UPDATE / MERGE, schema and partition evolution, partial V3 support (Row Lineage & Deletion Vector) | **Variant read and write** | — | | **Paimon** | Read only: Catalog access, incremental query, Time Travel, Branch/Tag, system tables, Deletion Vector reads | **Write support** (INSERT / UPDATE / DELETE / MERGE, DDL)<br>**Variant read and write**<br>Tracked in #65086 | Paimon index reads | | **Lance** | — | **Read support**: Catalog access, parallel scan, vector and full-text search, vector index management<br>Tracked in #66340 | — | | **Fluss** | — | — | **Read support**: log tables, primary-key tables, tiered tables (Paimon lake + log union read)<br>Proposed in #66399 | | **ADBC sources** | — | — | **ADBC Catalog**: access external sources over the standard ADBC and Arrow Flight SQL protocols<br>Proposed in #66331, tracked in #65615 | All three format tables below use the same capability dimensions so they can be compared side by side. An empty cell means the capability does not apply to that data source or is not currently planned — for example, Lance is read-only in 4.2, and Iceberg and Paimon do not offer vector search. ## Iceberg | Capability | Supported today (≤ 4.1) | 4.2 (2026-09) | 5.0 (2026-11) | |---|---|---|---| | **Catalog access** | ✅ HMS, REST, Hadoop, Glue, DLF, JDBC (4.1+), S3 Tables<br>✅ REST covers Apache Polaris, Apache Gravitino, Aliyun DLF 2.5+, AWS S3 Tables, AWS Glue REST<br>✅ Nested namespaces (3.1.2+, REST only) | — | — | | **Basic reads** | ✅ V1 / V2 table formats officially supported, V3 partially covered<br>✅ Parquet, ORC<br>✅ Position Delete, Equality Delete, Deletion Vector (4.1+) reads | — | — | | **Advanced reads** | ✅ Time Travel by snapshot or timestamp<br>✅ System tables (3.1+): snapshots, history, metadata_log_entries, entries, files, manifests, partitions, refs and others; all_manifests requires 4.0.4+<br>✅ Branch / Tag reads | — | — | | **Data writes** | ✅ `INSERT INTO`<br>✅ `INSERT OVERWRITE`, including branch writes (3.1+) and static/hybrid partition overwrite (4.1+)<br>✅ `CTAS` with file format, partitioning and table properties | — | — | | **Row-level updates** | 🔬 `DELETE`, `UPDATE`, `MERGE INTO` (4.1+, requires format-version ≥ 2), on both V2 and V3 tables | — | — | | **Table management** | ✅ Database and table DDL: CREATE / DROP DATABASE, CREATE / DROP TABLE<br>✅ Schema change (3.1+): add / drop / rename / modify / reorder top-level columns<br>✅ Limited complex-type changes (4.0.4+): safe type promotion, appending nullable struct fields<br>✅ Partition evolution (4.0.2+): add / drop / replace partition key, covering identity, bucket, truncate, year, month, day, hour<br>✅ Branch / Tag management (3.1+)<br>✅ `ORDER BY` at table creation to define sort order (4.1+)<br>🔬 Table actions: rollback, set_current_snapshot, cherrypick_snapshot, fast_forward, rewrite_data_files (4.0.2+); expire_snapshots, publish_changes, rewrite_manifests (4.1+) | — | — | | **Semi-structured (Variant)** | — | 🆕 Read and write Iceberg V3 Variant<br>🆕 Both plain and shredded layouts<br>🆕 Variant path extraction, CAST and predicate filtering | — | | **Multi-modal search** | — | — | — | | **Type mapping** | ✅ Common scalar types, STRUCT / MAP (2.1.3+)<br>✅ binary / uuid to varbinary (4.0.2+, `enable.mapping.varbinary`)<br>✅ timestamptz mapping (4.0.3+, `enable.mapping.timestamp_tz`) | 🆕 Variant | — | ## Paimon 4.2 turns the Paimon Catalog from read-only into a full read/write integration. The detailed feature checklist and current implementation status are tracked in #65086. | Capability | Supported today (≤ 4.1) | 4.2 (2026-09) | 5.0 (2026-11) | |---|---|---|---| | **Catalog access** | ✅ FileSystem, HMS, DLF<br>🔬 JDBC (4.1+)<br>✅ Built against Paimon 1.0.0 | — | 📅 Read the Paimon lake half of a tiered table through the Fluss Catalog | | **Basic reads** | ✅ Append tables, bucketed append tables, primary-key tables<br>✅ Parquet, ORC<br>✅ Deletion Vector reads<br>✅ Storage: HDFS, S3, GCS, OSS, COS, OBS, MinIO; Ozone requires 4.0.4+ | — | — | | **Advanced reads** | ✅ Batch incremental query (3.1+) by snapshot range or time range, with delta, changelog, diff and auto scan modes<br>✅ Time Travel (3.1+): `FOR TIME AS OF`, `FOR VERSION AS OF`; second-level and date formats from 3.1.2+<br>✅ Branch / Tag reads (3.1+): `@branch`, `@tag`, plus `$branches` and `$tags` metadata<br>✅ Table-level system tables (3.1+): schemas, snapshots, options, files, tags, branches, consumers, manifests, partitions, buckets, statistics, table_indexes and others | — | — | | **Data writes** | — | 🆕 `INSERT INTO ... SELECT` and `INSERT INTO ... VALUES`<br>🆕 `INSERT OVERWRITE`: non-partitioned, static partition, partial static partition, dynamic partition, and empty-input overwrite<br>🆕 Append-only tables and primary-key full-row writes<br>🆕 Merge engines: deduplicate, partial-update, aggregation, first-row<br>🆕 Bucket modes: fixed, unaware, postpone, hash dynamic, key dynamic<br>🆕 Primitive types, DECIMAL / DATE / TIMESTAMP, and ARRAY / MAP / STRUCT<br>🆕 Changelog producers: none, input, lookup, full-compaction<br>🆕 Atomic commit coordinated by FE, with idempotent retry, abort and rollback cleanup<br>🆕 FileSystem catalog, HDFS and object storage | — | | **Row-level updates** | — | 🆕 Row-level `UPDATE`, `DELETE` and `MERGE` | — | | **Table management** | ✅ `ALTER CATALOG ... SET PROPERTIES` to adjust Catalog parameters | 🆕 `CREATE TABLE` with table properties<br>🆕 `ADD` / `DROP` / `RENAME` / `MODIFY COLUMN` and column reorder<br>🆕 Schema change integrated with the write path | — | | **Semi-structured (Variant)** | — | 🆕 Read and write Paimon Variant<br>🆕 Both plain and shredded layouts<br>🆕 Covers append tables and primary-key tables<br>🆕 Variant path extraction, CAST and predicate filtering | — | | **Multi-modal search** | — | — | 📅 Read Paimon indexes directly | | **Type mapping** | ✅ Common scalar types plus array / map / row<br>✅ binary / varbinary map to string by default; 4.0.2+ can map to varbinary<br>✅ timestamp precision up to 6 digits<br>✅ timestamp with local time zone maps to datetime by default; 4.0.3+ can map to timestamptz<br>✅ `DESCRIBE` shows `WITH_TIMEZONE` (3.0.8+) | 🆕 Variant | — | ## Lance [Lance](https://lancedb.github.io/lance/) is a columnar format built for AI workloads, with native support for vectors, blobs and full-text indexes. Doris 4.2 adds Lance read support, bringing vector search and full-text search into the same SQL engine as the rest of the lakehouse. The full feature matrix, including capabilities explicitly out of scope for 4.2, is tracked in #66340. | Capability | Supported today (≤ 4.1) | 4.2 (2026-09) | 5.0 (2026-11) | |---|---|---|---| | **Catalog access** | — | 🆕 Filesystem Lance Catalog over a local path, `file://` or `s3://`<br>🆕 Lance REST Namespace Catalog<br>🆕 Authentication: bearer token, API key, custom HTTP headers<br>🆕 Temporary storage credentials issued by the namespace service<br>🆕 Multi-level Lance namespaces mapped to Doris databases | — | | **Basic reads** | — | 🆕 Query Lance datasets with parallel scan and column pruning<br>🆕 Scalar predicate pushdown; `vector_search()` accepts a SQL filter<br>🆕 `EXPLAIN` shows which predicates were pushed down<br>🆕 Read datasets directly through the `s3()` and `local()` TVFs, including glob patterns and multiple datasets in one call | — | | **Advanced reads** | — | 🆕 Time Travel and explicit version selection<br>🆕 The dataset version is pinned for the whole query, preserving snapshot consistency | — | | **Data writes** | — | — | — | | **Row-level updates** | — | — | — | | **Table management** | — | 🆕 `SHOW DATABASES`, `SHOW TABLES`, `DESC` | — | | **Semi-structured** | — | 🆕 JSON extension type reads | — | | **Multi-modal search** | — | 🆕 **Search modes**: ANN index search, flat exact search, radius/range search, BM25 full-text search, multi-vector search<br>🆕 **Vector index algorithms**: IVF_FLAT, IVF_SQ, IVF_PQ, IVF_HNSW_FLAT, IVF_HNSW_SQ, IVF_HNSW_PQ<br>🆕 **Distance metrics**: L2, cosine, dot product, hamming<br>🆕 **Index management**: create vector and scalar indexes, inspect index details, replace or rebuild an index under the same name, drop by name<br>🆕 **Index selection**: automatically select an index for the target vector column<br>🆕 **ANN tuning**: `nprobes`, `refine_factor`, `ef`, `use_index`<br>🆕 **Filtering**: scalar pre-filter and post-filter<br>🆕 **Result output**: Top-K, offset, and the `_distance` virtual column | — | | **Type mapping** | — | 🆕 Vector columns: `FixedSizeList<Float16 / Float32 / Float64 / UInt8 / Int8>`<br>🆕 Multi-vector columns: `List<FixedSizeList<...>>`<br>🆕 Blob v2, BFloat16 and other extension types | — | ## New Data Sources in 5.0 Where 4.2 goes deeper on Iceberg, Paimon and Lance, 5.0 goes wider — adding two classes of data source Doris cannot reach today. ### Fluss Catalog [Apache Fluss (incubating)](https://fluss.apache.org/) is a streaming storage layer built for real-time analytics. 5.0 adds a `fluss` Catalog that reads all three shapes a Fluss table can take: | Table shape | Read semantics | |---|---| | Log table | Reads the Fluss changelog | | Primary-key table | A KV snapshot plus the changelog written after it, merged by key | | Tiered table (`table.datalake.enabled`) | **The Paimon lake plus the log written after it**, merged in a single query | ```sql CREATE CATALOG fluss PROPERTIES ( "type" = "fluss", "fluss.bootstrap.servers" = "host:9123" ); SELECT * FROM fluss.db.tbl; -- lake + log, merged SELECT * FROM fluss.db.`tbl$lake`; -- the lake alone ``` The lake half of a tiered table reuses Doris's existing Paimon read path, so native ORC/Parquet reads, deletion vectors and the file cache all apply to it as well. Proposed in #66399. The lake catalog currently supports `filesystem` only; HMS and REST lake catalogs, predicate pushdown into the lake half, and the write path are follow-ups. ### ADBC Catalog 5.0 adds an `adbc` Catalog that reads external sources through an [Arrow Database Connectivity](https://arrow.apache.org/adbc/) driver. Compared with the JDBC Catalog, it offers two things the JDBC architecture cannot provide: **data is transferred in Arrow format, avoiding the per-value conversion a JDBC scan pays**, and **a single query can be read in parallel by several BE nodes**, each taking a different partition of the result — JDBC has no notion of result partitioning, so a scan is read by one BE. ```sql CREATE CATALOG remote_source PROPERTIES ( "type" = "adbc", "driver_url" = "libadbc_driver_flightsql.so", "uri" = "grpc://remote-doris:8070", "user" = "root", "password" = "" ); SELECT count(*) FROM remote_source.some_db.some_table; ``` | Capability | Description | |---|---| | Arrow-native transfer | Data arrives in Arrow format, avoiding JDBC's per-value conversion | | Parallel reads across BEs | One query is read in parallel by several BEs, each taking a different partition; controlled by `partitioned_read` and `max_partitions` | | Column pruning | Only the columns the query actually needs are requested from the source | | Predicate pushdown | `=`, `!=`, `<`, `<=`, `>`, `>=`, `IS [NOT] NULL`, `[NOT] IN` and their AND / OR / NOT combinations become the remote `WHERE`; pushdown affects performance only, never the rows returned | | LIMIT pushdown | Supported | | `COUNT(*)` optimization | Reads no column data from the source | | Metadata browsing | `SHOW DATABASES`, `SHOW TABLES`, `DESC`, `SHOW CREATE TABLE` and `information_schema` | | Type mapping | Automatic mapping to Doris types, including ARRAY / MAP / STRUCT, DECIMAL, dates, and timestamps with and without a time zone | | Metadata cache | 600 second TTL by default; any `REFRESH` statement clears it | | SQL dialect | Conservative ANSI SQL by default, overridable via `sql_dialect`; a doris dialect is provided | | General query surface | Joins against internal tables and other catalogs, aggregation, `ORDER BY`, `UNION`, subqueries, `SELECT ... INTO OUTFILE`, and MTMVs built on ADBC tables | Phase one targets Arrow Flight SQL sources and is the intended replacement for the `remote_doris` catalog. Adding another source requires only that driver's library, plus a dialect implementation where its SQL differs from ANSI. Proposed in #66331, tracked in #65615. The catalog is read-only; Doris ships no ADBC driver, so the library is placed by the operator. ## Related Issues and PRs | Item | Link | |---|---| | Paimon table write support | #65086 | | Lance feature matrix for Doris 4.2 | #66340 | | Arrow Flight SQL and ADBC Catalog | #65615 | | ADBC Catalog implementation | #66331 | | Fluss Catalog implementation | #66399 | ## Feedback If your workload depends on something not covered above, or on a capability currently scheduled later than you need, please comment. Priorities for 4.2 and 5.0 are still being adjusted, and concrete use cases are the most useful input for that. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
