This is an automated email from the ASF dual-hosted git repository.

imbajin pushed a commit to branch add-seatunnel-integration-doc
in repository https://gitbox.apache.org/repos/asf/hugegraph-doc.git

commit 3c42928c8988ad5015d696d60db125f8c36ff83d
Author: dark <[email protected]>
AuthorDate: Thu Aug 20 00:47:38 2026 +0800

    docs(seatunnel): refine connector quick start
    
    - clarify Tools, Loader, and SeaTunnel selection\n- align JDBC, Kafka, and 
dev migration examples\n- add concise diagrams and fold long configs\n- replace 
stale architecture asset and version links
---
 .../seatunnel/hugegraph-seatunnel-architecture.png | Bin 1056077 -> 0 bytes
 .../images/seatunnel/seatunnel-graph2graph.png     | Bin 0 -> 825977 bytes
 .../images/seatunnel/seatunnel-kafka2graph.png     | Bin 0 -> 832987 bytes
 .../docs/images/seatunnel/seatunnel-overview.png   | Bin 0 -> 881633 bytes
 .../docs/images/seatunnel/seatunnel-sql2graph.png  | Bin 0 -> 875187 bytes
 content/cn/docs/introduction/_index.md             |  17 +-
 .../toolchain/hugegraph-seatunnel-connector.md     | 392 +++++++++------------
 7 files changed, 178 insertions(+), 231 deletions(-)

diff --git 
a/content/cn/docs/images/seatunnel/hugegraph-seatunnel-architecture.png 
b/content/cn/docs/images/seatunnel/hugegraph-seatunnel-architecture.png
deleted file mode 100644
index bf73bcb68..000000000
Binary files 
a/content/cn/docs/images/seatunnel/hugegraph-seatunnel-architecture.png and 
/dev/null differ
diff --git a/content/cn/docs/images/seatunnel/seatunnel-graph2graph.png 
b/content/cn/docs/images/seatunnel/seatunnel-graph2graph.png
new file mode 100644
index 000000000..5b0033df6
Binary files /dev/null and 
b/content/cn/docs/images/seatunnel/seatunnel-graph2graph.png differ
diff --git a/content/cn/docs/images/seatunnel/seatunnel-kafka2graph.png 
b/content/cn/docs/images/seatunnel/seatunnel-kafka2graph.png
new file mode 100644
index 000000000..96a0bed02
Binary files /dev/null and 
b/content/cn/docs/images/seatunnel/seatunnel-kafka2graph.png differ
diff --git a/content/cn/docs/images/seatunnel/seatunnel-overview.png 
b/content/cn/docs/images/seatunnel/seatunnel-overview.png
new file mode 100644
index 000000000..4bd0d2c39
Binary files /dev/null and 
b/content/cn/docs/images/seatunnel/seatunnel-overview.png differ
diff --git a/content/cn/docs/images/seatunnel/seatunnel-sql2graph.png 
b/content/cn/docs/images/seatunnel/seatunnel-sql2graph.png
new file mode 100644
index 000000000..71e4cf9a8
Binary files /dev/null and 
b/content/cn/docs/images/seatunnel/seatunnel-sql2graph.png differ
diff --git a/content/cn/docs/introduction/_index.md 
b/content/cn/docs/introduction/_index.md
index 8de7a852b..94e3debe3 100644
--- a/content/cn/docs/introduction/_index.md
+++ b/content/cn/docs/introduction/_index.md
@@ -23,18 +23,9 @@ HugeGraph 支持百亿以上的顶点和边的快速存储与查询,具备出
 
 ### 生态系统全景
 
-```text
-┌────────────────────────────────────────────────────────────────────┐
-│           Apache HugeGraph - Full-Stack Graph System               │
-├──────────────────┬────────────────────┬────────────────────────────┤
-│  Graph DB (OLTP) │    Graph Compute   │         Graph AI           │
-│  HugeGraph       │  Vermeer (Memory)  │      HugeGraph-AI          │
-│  Server          │  Computer (Dist.)  │    GraphRAG/GNN/Py         │
-├──────────────────┴────────────────────┴────────────────────────────┤
-│                      HugeGraph Toolchain                           │
-│  Hubble | Loader | Client(Java/Go/Py) | Spark | SeaTunnel | Tools  │
-└────────────────────────────────────────────────────────────────────┘
-```
+![HugeGraph 生态与 SeaTunnel 
数据流](/cn/docs/images/seatunnel/seatunnel-overview.png)
+
+HugeGraph 负责图存储和查询,Toolchain 负责导入、管理和连接外部数据系统。需要把 JDBC、Kafka 等数据接入图数据库时,可以从 
[SeaTunnel Connector 
快速开始](/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector) 开始。
 
 ---
 
@@ -94,7 +85,7 @@ HugeGraph 独立的 AI 组件,连接图与大语言模型(LLM):
 | [Loader](/cn/docs/quickstart/toolchain/hugegraph-loader) | 
数据导入工具:支持本地文件、HDFS、MySQL 等多数据源,TXT/CSV/JSON 等格式 |
 | [Client](/cn/docs/quickstart/client/hugegraph-client) | 多语言 SDK:Java / 
Python / Go |
 | [Spark-connector](/cn/docs/quickstart/toolchain/hugegraph-spark-connector) | 
Spark 集成:支持通过 Spark 批量读写图数据,适合大数据离线处理场景 |
-| [SeaTunnel 
Connector](/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector) | 提供 
HugeGraph Source 与 Sink,支持读取和写入图数据 |
+| [SeaTunnel 
Connector](/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector) | 提供 
HugeGraph Sink;Source 目前随 SeaTunnel dev 分支预览 |
 | [Tools](/cn/docs/quickstart/toolchain/hugegraph-tools) | 
命令行运维工具:图管理、备份恢复、Gremlin 执行等 |
 
 ---
diff --git 
a/content/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector.md 
b/content/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector.md
index 42323b920..bc423ec8f 100644
--- a/content/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector.md
+++ b/content/cn/docs/quickstart/toolchain/hugegraph-seatunnel-connector.md
@@ -4,126 +4,101 @@ linkTitle: "使用 SeaTunnel Connector 同步数据"
 weight: 5
 ---
 
-<!--
-  TODO(apache/hugegraph-doc#464):本文 SeaTunnel 官方文档链接暂指向 latest(无版本号)页面。
-  2.3.13 及之前版本的参数以官网版本下拉里的 2.3.13 文档为准。
-  HugeGraph Source Connector 尚未随正式版本发布,官网 latest 无对应页面,
-  故 Source 文档链接暂指向 GitHub dev 分支文件。
-  待 Source 随正式版本发布后,将链接切换为官网版本化地址,
-  例如 https://seatunnel.apache.org/docs/2.3.14/connectors/source/HugeGraph/
--->
+SeaTunnel 负责连接数据源和数据目的地。HugeGraph Connector 提供两种能力。
 
-### 1 版本与兼容性
+- `HugeGraph Sink` 把文件、数据库、Kafka 等数据写入 HugeGraph。
+- `HugeGraph Source` 从 HugeGraph 读出顶点和边,目前只在 SeaTunnel `dev` 分支提供。
 
-| 组件 | 版本要求 | 说明 |
-|------|---------|------|
-| Java | 11+ | HugeGraph Client 1.5.0+ 运行环境要求 |
-| Apache SeaTunnel | 2.3.13+ | - |
-| HugeGraph Server | 1.5.0+ | 需与 Connector 内置 Client 版本匹配 |
+![SeaTunnel 与 HugeGraph 
数据流总览](/cn/docs/images/seatunnel/seatunnel-overview.png)
 
-> 配置 API 有两代,别混用:
->
-> - **2.3.13 正式版**:Sink 只支持 `schema_config`(必填)。本文第 4 节和第 5.1 节的示例按它来写。
-> - **next(dev)分支**:`mappings` 多映射、`schema_save_mode` 自动建 
Schema、`data_save_mode` 等新选项,以及 **Source Connector** 都只在这里,正式版还没带。本文 5.3 节的图迁移是 
dev 预览。
->
-> next 分支内置的 HugeGraph Client 已升到 1.7.0,需搭配对应版本的 Server。
+## 1 先看版本
 
-### 2 概述
+本文把发布版和开发版分开写。配置放错版本,任务会在启动阶段失败。
 
-[Apache SeaTunnel](https://seatunnel.apache.org/) 是开源数据集成平台,批处理、流式同步都支持,自带 
100+ 连接器。HugeGraph 的 Connector-V2 已合入:Sink 负责写入(顶点/边的批量写入、更新、删除),Source 
负责读取(批量读取、整图迁移,尚未发布)。
+| 使用内容 | SeaTunnel 版本 | 配置方式 | 状态 |
+| --- | --- | --- | --- |
+| JDBC / Kafka 写入 HugeGraph | 2.3.13 | `schema_config` | 发布版 |
+| HugeGraph 读取和图迁移 | 当前 `dev` | HugeGraph Source + `mappings` | 开发预览 |
+| `graph2graph` | 当前 `dev` | Source + Sink | 开发预览 |
 
-![HugeGraph + SeaTunnel 
数据集成架构图](/cn/docs/images/seatunnel/hugegraph-seatunnel-architecture.png)
+当前 `dev` 示例按 commit 
[`f1a1a0a`](https://github.com/apache/seatunnel/commit/f1a1a0abbe24bdac8cf23307995a78f778a3f467)
 核对,固定版本的 [HugeGraph Source 
文档](https://github.com/apache/seatunnel/blob/f1a1a0abbe24bdac8cf23307995a78f778a3f467/docs/zh/connectors/source/HugeGraph.md)
 和 [HugeGraph Sink 
文档](https://github.com/apache/seatunnel/blob/f1a1a0abbe24bdac8cf23307995a78f778a3f467/docs/zh/connectors/sink/HugeGraph.md)
 与本文对应。`dev` 会继续变化,使用新版本前请重新核对配置。
 
-### 3 选型:SeaTunnel 还是 Loader
+2.3.13 的 HugeGraph Sink 使用 `schema_config`,这个版本没有 HugeGraph 
Source。本文所有发布版示例都按这个边界编写。
 
-可以把 SeaTunnel 理解为 [Loader](/cn/docs/quickstart/toolchain/hugegraph-loader) + 
[Tools](/cn/docs/quickstart/toolchain/hugegraph-tools) 的合集。两者不冲突:Loader 只管 
HugeGraph,开箱即用;SeaTunnel 面向所有数据系统。按场景选:
+## 2 选哪个工具
 
-| 你的场景 | 推荐 | 理由 |
-|---------|------|------|
-| 一次性 / 定时把本地文件、HDFS、MySQL 等导入 HugeGraph | Loader | 免部署,映射文件即写即用 |
-| 数据已在(或必须经过)Flink、Spark、Kafka 等大数据管道 | SeaTunnel | 复用现有管道,不引入第二套导入工具 |
-| 实时流式导入、CDC 增量同步 | SeaTunnel | 原生流式作业 + checkpoint 断点恢复 |
-| 只维护 HugeGraph 一张图,规模可控、追求简单 | Loader | 工具链内闭环,无额外集群 |
+先看数据从哪里来,以及任务是否已经属于一条大数据管道。
 
-选 SeaTunnel 要接受两点:部署一套 SeaTunnel 集群(或复用现有的),作业用通用 HOCON 配置,而不是 HugeGraph 映射文件。
+| 你的任务 | 推荐工具 | 适合原因 |
+| --- | --- | --- |
+| 管理图、执行 Gremlin、备份恢复、图克隆 | 
[HugeGraph-Tools](/cn/docs/quickstart/toolchain/hugegraph-tools) | 只操作 
HugeGraph,命令直接 |
+| 把本地文件、HDFS、MySQL 等数据批量导入 HugeGraph | 
[HugeGraph-Loader](/cn/docs/quickstart/toolchain/hugegraph-loader) | 配置简单,导入流程短 
|
+| 数据要经过 Kafka、JDBC、Flink、Spark 或多个外部系统 | SeaTunnel | 可以复用已有数据管道 |
+| 需要流式任务、checkpoint 或统一管理多个连接器 | SeaTunnel | 支持 Source、Transform 和 Sink 组合 |
+| 稳定地把一张 HugeGraph 图复制到另一张图 | Tools 优先 | 发布版工具更直接;SeaTunnel Source 仍是 dev 预览 |
 
-### 4 快速开始
+只维护一张图、没有现成大数据管道时,优先从 Loader 或 Tools 开始。SeaTunnel 需要额外准备连接器插件,并使用 HOCON 配置文件。
 
-前置条件:HugeGraph Server 
1.5.0+([部署指南](/cn/docs/quickstart/hugegraph/hugegraph-server))。
+## 3 准备工作
 
-#### 4.1 部署 SeaTunnel
+### 3.1 HugeGraph
 
-**Docker(推荐)**。注意两点:官方 `apache/seatunnel:2.3.13` 镜像只内置 fake/console 
两个连接器([官方说明](https://seatunnel.apache.org/docs/getting-started/docker/)),且基础镜像是 
JDK8;本文示例需要 LocalFile、HugeGraph 连接器和 Java 11。所以官方镜像不能直接跑本文作业,二选一:
+本文示例使用以下图模型。
 
-- 基于 JDK11 自建镜像并装好插件(示意,细节以官方 
[自建镜像文档](https://seatunnel.apache.org/docs/getting-started/docker/) 为准):
+| 图元素 | 配置 |
+| --- | --- |
+| VertexLabel | `person`,主键为 `name` |
+| PropertyKey | `name` 为 Text,`age` 为 Int |
+| EdgeLabel | `knows`,源和目标都是 `person`,属性为 `since` |
 
-```dockerfile
-FROM eclipse-temurin:11-jre
-RUN curl -L -o /tmp/st.tgz 
https://downloads.apache.org/seatunnel/2.3.13/apache-seatunnel-2.3.13-bin.tar.gz
 \
-    && tar -xzf /tmp/st.tgz -C /opt \
-    && mv /opt/apache-seatunnel-2.3.13 /opt/seatunnel \
-    && sh /opt/seatunnel/bin/install-plugin.sh 2.3.13 \
-    && rm /tmp/st.tgz
-WORKDIR /opt/seatunnel
-```
-
-```bash
-docker build -t seatunnel-hg:2.3.13 .
-```
+2.3.13 的 Sink 会按 `schema_config` 读取已有的 VertexLabel、EdgeLabel 和 
PropertyKey。运行写入任务前,请先在 Hubble、REST API 或 Gremlin 中创建 Schema。
 
-- 或者直接用 4.1 末的二进制包方式跑(最省事)。
+### 3.2 SeaTunnel
 
-**Kubernetes**:生产集群用 Helm 部署,见官方 
[K8s(Helm)部署文档](https://seatunnel.apache.org/docs/getting-started/kubernetes/helm/)。
+请按 [SeaTunnel 
本地部署文档](https://seatunnel.apache.org/docs/getting-started/locally/deployment/) 
获取发行包。2.2.0-beta 之后,发行包默认不带连接器依赖,需要按任务安装 JDBC、Kafka 和 HugeGraph 插件;JDBC 
还需要对应数据库的驱动。
 
-**二进制包(参考)**:从 [下载页](https://seatunnel.apache.org/download/) 
取安装包,解压后装插件、跑脚本。JVM 脚本方式仅作本地调试参考,生产优先 Docker / K8s:
+如果 SeaTunnel 与 HugeGraph 不在同一台机器,`host` 要填写 SeaTunnel 运行环境可以访问的地址。容器内的 
`127.0.0.1` 指向 SeaTunnel 容器自身;同一 Docker 网络中的服务则使用 HugeGraph 的服务名。
 
-```bash
-sh bin/install-plugin.sh 2.3.13
-sh bin/seatunnel.sh --config ./config/hugegraph-sync.conf -e local
-```
+## 4 sql2graph
 
-部署细节以官方 
[本地部署文档](https://seatunnel.apache.org/docs/getting-started/locally/deployment/) 
为准。
+JDBC 方式适合把关系库中的表或 SQL 查询结果导入 HugeGraph。下面的例子把 `person` 表写成顶点,使用 `name` 生成 
HugeGraph 主键。
 
-#### 4.2 最小示例:CSV 文件 → person 顶点
+![关系表写入 HugeGraph](/cn/docs/images/seatunnel/seatunnel-sql2graph.png)
 
-先建 Schema:PropertyKey `name`(Text)、`age`(Int),VertexLabel `person`(主键 
`name`),边示例另需 EdgeLabel `knows`(属性 `since`)。2.3.13 的 Sink 不会自动建 Schema,可在 
Hubble 里建,或用服务端 REST/Gremlin 建。
+### 4.1 关系库到顶点
 
-准备一个无表头的 CSV(列顺序与 `schema.fields` 声明顺序一致),内容如下:
+假设 MySQL 中有一张表。
 
-```csv
-marko,29
-vadas,27
-josh,32
+```sql
+CREATE TABLE person (
+  name VARCHAR(64) PRIMARY KEY,
+  age INT NOT NULL
+);
 ```
 
+在 SeaTunnel 安装目录下创建 `config/sql2graph-person.conf`。
+
 ```hocon
 env {
   job.mode = "BATCH"
 }
 
 source {
-  LocalFile {
-    path = "/data/person.csv"
-    file_format_type = "csv"
-    schema = {
-      fields = {
-        name = "string"
-        age = "int"
-      }
-    }
+  Jdbc {
+    url = "jdbc:mysql://mysql:3306/demo?useSSL=false&serverTimezone=UTC"
+    driver = "com.mysql.cj.jdbc.Driver"
+    username = "seatunnel"
+    password = "change_me"
+    query = "SELECT name, age FROM person ORDER BY name"
   }
 }
 
 sink {
   HugeGraph {
-    host = "127.0.0.1"
+    host = "hugegraph"
     port = 8080
     graph_name = "hugegraph"
-    graph_space = "DEFAULT"
-    # 以下为可选参数,默认值见第 6 节;认证开启时再填 username/password
-    # protocol = "http"
-    # username = "admin"
-    # password = "admin"
+    graph_space = "default"
     schema_config = {
       type = "VERTEX"
       label = "person"
@@ -135,98 +110,72 @@ sink {
 }
 ```
 
-一行 CSV 到图顶点的映射过程:
-
-```text
-┌───────────────── 输入行 (SeaTunnel Row) ─────────────────┐
-│  name = "marko"        age = 29                          │
-└──────────────────────────┬───────────────────────────────┘
-                           │ schema_config:
-                           │   type = VERTEX, label = person
-                           │   idStrategy = PRIMARY_KEY, idFields = [name]
-                           ▼
-                 ┌───────────────────────────────┐
-                 │ HugeGraph 顶点                 │
-                 │ id   = person:marko           │
-                 │ label = person                │
-                 │ props = { name: "marko",      │
-                 │           age: 29 }           │
-                 └───────────────────────────────┘
-```
-
-#### 4.3 运行与验证
+执行任务。
 
 ```bash
-# 二进制方式(4.1 的自建镜像同理,把挂载和容器名换一下)
-sh bin/seatunnel.sh --config ./config/hugegraph-sync.conf -e local
+./bin/seatunnel.sh --config ./config/sql2graph-person.conf -m local
 ```
 
-> **容器里跑要注意网络**:`127.0.0.1` 指向容器自身,连不到宿主机或另一个容器。
->
-> - HugeGraph 跑在宿主机:把 `host` 改成 `host.docker.internal`,Linux 启动容器时加 
`--add-host=host.docker.internal:host-gateway`;
-> - HugeGraph 也在容器里:两个容器进同一个 docker 网络,`host` 填 HugeGraph 容器的容器名/服务名。
-
-运行后,在 Hubble 或 REST API 里执行这条 Gremlin 验证:
+执行后可以在 HugeGraph 中检查顶点。
 
 ```groovy
-g.V().hasLabel('person').valueMap()
+g.V().hasLabel('person').valueMap('name', 'age')
 ```
 
-#### 4.4 写入边
+`Jdbc` 的 `url` 和 `driver` 必填。`username` 和 `password` 
按数据库认证配置填写,匿名连接时可以省略;示例中的密码需要替换。MySQL 驱动需要放到 SeaTunnel 对应引擎的插件目录,具体位置见 [JDBC 
Source 文档](https://seatunnel.apache.org/docs/connectors/source/Jdbc/)。
 
-关系行到边的映射同样用 `schema_config`,`sourceConfig` / `targetConfig` 还原边的两个端点,字段改名放在 
`mapping.fieldMapping` 里。
-
-<details>
-<summary>展开查看:关系 CSV → knows 边 完整配置</summary>
+### 4.2 关系库到边
 
-CSV(无表头):
+如果关系表中的端点字段已经能直接对应 `person.name`,可以再运行一个边任务。假设表结构如下。
 
-```csv
-marko,vadas,2020
-marko,josh,2021
+```sql
+CREATE TABLE knows (
+  source_name VARCHAR(64) NOT NULL,
+  target_name VARCHAR(64) NOT NULL,
+  since INT NOT NULL
+);
 ```
 
+<details>
+<summary>展开查看边任务配置</summary>
+
 ```hocon
 env {
   job.mode = "BATCH"
 }
 
 source {
-  LocalFile {
-    path = "/data/knows.csv"
-    file_format_type = "csv"
-    schema = {
-      fields = {
-        person1_name = "string"
-        person2_name = "string"
-        since = "int"
-      }
-    }
+  Jdbc {
+    url = "jdbc:mysql://mysql:3306/demo?useSSL=false&serverTimezone=UTC"
+    driver = "com.mysql.cj.jdbc.Driver"
+    username = "seatunnel"
+    password = "change_me"
+    query = "SELECT source_name, target_name, since FROM knows ORDER BY 
source_name, target_name"
   }
 }
 
 sink {
   HugeGraph {
-    host = "127.0.0.1"
+    host = "hugegraph"
     port = 8080
     graph_name = "hugegraph"
-    graph_space = "DEFAULT"
+    graph_space = "default"
     schema_config = {
       type = "EDGE"
       label = "knows"
       sourceConfig = {
         label = "person"
-        idFields = ["person1_name"]
+        idFields = ["source_name"]
       }
       targetConfig = {
         label = "person"
-        idFields = ["person2_name"]
+        idFields = ["target_name"]
       }
       properties = ["since"]
       mapping = {
         fieldMapping = {
-          person1_name = "name"
-          person2_name = "name"
+          source_name = "name"
+          target_name = "name"
         }
       }
     }
@@ -236,34 +185,23 @@ sink {
 
 </details>
 
-```text
-person1_name = "marko"   person2_name = "vadas"   since = 2020
-        │                        │
-        ▼                        ▼
-  sourceConfig             targetConfig
-  label = person           label = person
-  idFields = [person1_name] idFields = [person2_name]
-        └───────────┬────────────┘
-                    ▼
-      edge: person:marko -[knows]-> person:vadas
-      properties = { since: 2020 }
-```
+先写顶点,再写边。端点字段如果只是外键,不能直接拼出 HugeGraph 顶点 ID,需要先在 SQL 中完成关联查询,或者先把端点名称写入结果集。
+
+CDC 配置请参考 SeaTunnel 的 [MySQL CDC 
文档](https://seatunnel.apache.org/docs/connectors/source/MySQL-CDC/)。
+
+## 5 kafka2graph
 
-> 写边时端点顶点还不存在的话,默认行为可能产生孤儿边或幻影顶点(服务端行为)。不能接受就先导顶点、再导边;所用版本支持 `check_vertex` 
时,开 `check_vertex = true` 可让服务端直接拒绝这类边(以所用版本的官方文档为准)。
+Kafka 适合持续把事件写入 HugeGraph。下面的消息使用 JSON 格式,每条消息对应一个 `person` 顶点。
 
-### 5 常见场景
+![Kafka 事件写入 HugeGraph](/cn/docs/images/seatunnel/seatunnel-kafka2graph.png)
 
-#### 5.1 Kafka 实时流导入
+Kafka topic `user-events` 中的消息示例。
 
-```mermaid
-flowchart LR
-    P["业务系统"] --> K[("Kafka Topic")]
-    K -->|"STREAMING + checkpoint"| ST["SeaTunnel<br/>Kafka Source → HugeGraph 
Sink"]
-    ST --> HG[("HugeGraph Server")]
+```json
+{"name":"marko","age":29}
 ```
 
-<details>
-<summary>展开查看:Kafka → HugeGraph 流式作业配置(2.3.13)</summary>
+创建 `config/kafka2graph.conf`。
 
 ```hocon
 env {
@@ -273,10 +211,11 @@ env {
 
 source {
   Kafka {
-    bootstrap.servers = "localhost:9092"
+    bootstrap.servers = "kafka:9092"
     topic = "user-events"
     consumer.group = "hugegraph-import"
-    format = json
+    start_mode = "earliest"
+    format = "json"
     schema = {
       fields = {
         name = "string"
@@ -288,10 +227,10 @@ source {
 
 sink {
   HugeGraph {
-    host = "127.0.0.1"
+    host = "hugegraph"
     port = 8080
     graph_name = "hugegraph"
-    graph_space = "DEFAULT"
+    graph_space = "default"
     schema_config = {
       type = "VERTEX"
       label = "person"
@@ -303,26 +242,33 @@ sink {
 }
 ```
 
-</details>
+```bash
+./bin/seatunnel.sh --config ./config/kafka2graph.conf -m local
+```
 
-> 要点:`job.mode = "STREAMING"` 加 `checkpoint.interval`,任务就能断点恢复。Sink 是 
at-least-once 语义,`PRIMARY_KEY` / `CUSTOMIZE_*` 这类可还原的 ID 重放只是幂等更新。Kafka 参数是 
`topic`(逗号分隔多 topic),offset、分区、format 等见 [Kafka Source 
官方文档](https://seatunnel.apache.org/docs/connectors/source/Kafka/)。
+`checkpoint.interval` 用于保存任务状态。HugeGraph Sink 使用 at-least-once 写入语义,使用 
`PRIMARY_KEY` 时,重复写入同一个 `name` 会落到同一个顶点,不会因为重放生成新的随机顶点 ID。
 
-#### 5.2 在 Flink / Spark 上运行
+Kafka 的参数和消息格式见 [Kafka Source 
文档](https://seatunnel.apache.org/docs/connectors/source/Kafka/)。写边时,把 Sink 的 
`schema_config.type` 改为 `EDGE`,再补充 `sourceConfig`、`targetConfig` 和边属性。
 
-作业默认跑在 SeaTunnel 自带的 Zeta 引擎上。已有 Flink / Spark 集群的话,同一份配置换对应的 starter 
脚本提交即可,内容不用改。脚本用法和版本支持见官方 [Flink 
引擎文档](https://seatunnel.apache.org/docs/engines/flink/) 与 [Spark 
引擎文档](https://seatunnel.apache.org/docs/engines/spark/)。
+## 6 graph2graph
 
-#### 5.3 HugeGraph → HugeGraph 图迁移(dev 预览)
+HugeGraph Source 目前只在 SeaTunnel `dev` 分支提供。下面的配置按 
[`f1a1a0a`](https://github.com/apache/seatunnel/tree/f1a1a0abbe24bdac8cf23307995a78f778a3f467)
 核对,不能直接放进 2.3.13 发行包。
 
-2.3.13 没有 HugeGraph Source,本节需要 next(dev)构建(或等正式发布);dev 同时支持 `mappings` 新配置。
+![HugeGraph 图迁移](/cn/docs/images/seatunnel/seatunnel-graph2graph.png)
 
-```mermaid
-flowchart LR
-    A[("HugeGraph 源实例")] -->|"Source 批量读取"| ST["SeaTunnel"]
-    ST -->|"Sink 批量写入"| B[("HugeGraph 目标实例")]
-```
+`mappings` 默认会创建缺失的 Schema。边映射的源和目标顶点标签仍需存在,因此要按下面的顺序先跑顶点任务,再跑边任务;如果把 
`schema_save_mode` 改成 `ERROR_WHEN_SCHEMA_NOT_EXIST`,请提前创建目标图 Schema。
+
+一次迁移按两个任务执行。
+
+1. 先迁移顶点。
+2. 再迁移边。
+
+### 6.1 迁移顶点
+
+下面的 Source 读取源图的 `person` 顶点,Sink 使用 `name` 重新生成 `PRIMARY_KEY` 顶点 ID。
 
 <details>
-<summary>展开查看:单图迁移(person 顶点,dev 构建)</summary>
+<summary>展开查看顶点迁移配置</summary>
 
 ```hocon
 env {
@@ -331,7 +277,7 @@ env {
 
 source {
   HugeGraph {
-    host = "graph-a"
+    host = "source-hugegraph"
     port = 8080
     graph_name = "hugegraph"
     graph_space = "DEFAULT"
@@ -348,7 +294,7 @@ source {
 
 sink {
   HugeGraph {
-    host = "graph-b"
+    host = "target-hugegraph"
     port = 8080
     graph_name = "hugegraph"
     graph_space = "DEFAULT"
@@ -367,80 +313,90 @@ sink {
 
 </details>
 
-<details>
-<summary>展开查看:批量迁移多个图的思路(示意,非可直接运行的配置)</summary>
+### 6.2 迁移边
 
-A 实例有 3 个图,全部迁到 B 实例、图名不变时,按下面的思路组织,而不是复制粘贴下面这段伪代码:
+Source 会为边补充 `~source_id` 和 `~target_id` 保留列。Sink 可以直接使用这两列还原端点 ID。
 
-1. **一次作业只迁一种元素**:Source 的 `label_type` 只能是 `VERTEX` 或 `EDGE` 
之一,顶点、边各跑一次;先顶点、后边(边依赖端点)。
-2. **省略 `label` 读全量**:Source 省略 `label` 会按 label 输出多张输入表;此时 sink 的 `mappings` 
必须按每个 label 逐条写全,并按官方文档把每条映射绑定到对应输入表,否则会交叉写入。单 label 的 mappings 不能直接复用。
-3. **多图用变量替换**:配置里写 `graph_name = "${graph}"`,提交时 `-i graph=xxx`(多个参数用逗号分隔,见官方 
[命令文档](https://seatunnel.apache.org/docs/engines/zeta/user-command/))。
+<details>
+<summary>展开查看边迁移配置</summary>
 
 ```hocon
-# 伪代码:结构示意,mappings 需按 label 逐条补全并绑定输入表
 env {
   job.mode = "BATCH"
 }
 
 source {
   HugeGraph {
-    host = "graph-a"
+    host = "source-hugegraph"
     port = 8080
-    graph_name = "${graph}"
+    graph_name = "hugegraph"
     graph_space = "DEFAULT"
-    label_type = "VERTEX"
-    # 不写 label:读取该图全部顶点 label;迁边时改成 "EDGE"
+    label = "knows"
+    label_type = "EDGE"
+    schema = {
+      fields = {
+        since = "int"
+      }
+    }
   }
 }
 
 sink {
   HugeGraph {
-    host = "graph-b"
+    host = "target-hugegraph"
     port = 8080
-    graph_name = "${graph}"
+    graph_name = "hugegraph"
     graph_space = "DEFAULT"
-    mappings = [ /* 按 label 逐条写,并绑定对应输入表 */ ]
+    mappings = [
+      {
+        type = "EDGE"
+        label = "knows"
+        sourceConfig = {
+          label = "person"
+          idFields = ["~source_id"]
+        }
+        targetConfig = {
+          label = "person"
+          idFields = ["~target_id"]
+        }
+        properties = ["since"]
+      }
+    ]
   }
 }
 ```
 
-```bash
-# 每张图:先跑 VERTEX 作业,再改 label_type 跑 EDGE 作业
-for g in graph1 graph2 graph3; do
-  sh bin/seatunnel.sh --config ./config/hugegraph-migrate.conf -e local -i 
graph=$g
-done
-```
-
 </details>
 
-> 克隆边时,Source 输出自带保留列 `~source_id` / `~target_id`,Sink 直接配 
`sourceConfig.idFields = ["~source_id"]`、`targetConfig.idFields = 
["~target_id"]` 就能复用,不用重新拼 ID。Source 还支持 `parallelism > 1` 分片并行,详见 [HugeGraph 
Source 
文档](https://github.com/apache/seatunnel/blob/dev/docs/zh/connectors/source/HugeGraph.md)。
+如果源图使用 `AUTOMATIC` 顶点 ID,Source 无法把原 ID 作为主键重新生成。需要保留原 ID 时,应按 dev 文档使用 
`CUSTOMIZE_*` 策略和 `~id` 保留列,并先确认目标图 Schema 与 ID 策略一致。
 
-### 6 核心配置参数
+## 7 常用配置
 
-以下为 2.3.13 版本 Sink 的常用参数;`mappings`、`schema_save_mode`、`data_save_mode` 等是 
next(dev)分支新增(5.3 节用到 `mappings`),完整说明以官网对应版本文档为准:
+以下字段同时出现在本文的发布版示例中。
 
-| 参数 | 类型 | 必填 | 默认值 | 说明 |
-|------|------|------|--------|------|
-| `host` | String | 是 | - | HugeGraph Server 地址(只填主机名/IP,端口走 `port`) |
-| `port` | Integer | 是 | - | HugeGraph Server 端口 |
-| `graph_name` | String | 是 | - | 图名称 |
-| `schema_config` | Object | 是 | - | 顶点/边映射(2.3.13 的配置方式) |
-| `graph_space` | String | 否 | `DEFAULT` | 图空间 |
-| `protocol` | String | 否 | `http` | 服务协议,支持 `http` / `https` |
-| `username` | String | 否 | - | 认证用户名(服务端开启认证时必填) |
-| `password` | String | 否 | - | 认证密码(服务端开启认证时必填) |
-| `batch_size` | Integer | 否 | 500 | 单批写入前缓冲的记录数 |
-| `batch_interval_ms` | Integer | 否 | 5000 | 刷新批次的最大等待时间(毫秒) |
+| 字段 | 作用 |
+| --- | --- |
+| `host` | HugeGraph Server 主机名或 IP,不要把端口写进来 |
+| `port` | HugeGraph Server 端口 |
+| `graph_name` | 图名称 |
+| `graph_space` | 图空间,发布版示例沿用 2.3.13 官方配置的 `default`,dev 示例使用源码默认值 
`DEFAULT`,不要跨版本复制 |
+| `schema_config` | 2.3.13 Sink 的顶点或边映射 |
+| `mappings` | dev Sink 的多映射配置 |
+| `batch_size` | 单批写入的记录数,默认值为 500 |
+| `batch_interval_ms` | 批次刷新等待时间,默认值为 5000 毫秒 |
 
-`schema_config` 的常用字段:`type`(`VERTEX` / 
`EDGE`)、`label`、`properties`、`idStrategy` / `idFields`(顶点)、`sourceConfig` / 
`targetConfig`(边)、`mapping.fieldMapping`。其余字段见官方文档。
+2.3.13 不支持本文 dev 示例中的 `mappings`、HugeGraph Source 和 
`schema_save_mode`。遇到配置校验失败时,先检查 SeaTunnel 发行包版本和配置 API 是否对应。
 
-### 参考文档
+## 8 参考文档
 
-- [Apache SeaTunnel 官方网站](https://seatunnel.apache.org/)
-- [HugeGraph Sink Connector 
官方文档](https://seatunnel.apache.org/docs/connectors/sink/HugeGraph/)(latest 展示 
`mappings` 新 API,2.3.13 的 `schema_config` 版从官网版本下拉切换)
-- [HugeGraph Source Connector 文档(dev 
分支)](https://github.com/apache/seatunnel/blob/dev/docs/zh/connectors/source/HugeGraph.md)
-- [SeaTunnel Docker 
部署](https://seatunnel.apache.org/docs/getting-started/docker/)
-- [SeaTunnel 
K8s(Helm)部署](https://seatunnel.apache.org/docs/getting-started/kubernetes/helm/)
 - [SeaTunnel 
本地部署](https://seatunnel.apache.org/docs/getting-started/locally/deployment/)
+- [HugeGraph Sink 
2.3.13](https://github.com/apache/seatunnel/blob/2.3.13/docs/zh/connectors/sink/HugeGraph.md)
+- [HugeGraph Sink 
dev(本文核对版本)](https://github.com/apache/seatunnel/blob/f1a1a0abbe24bdac8cf23307995a78f778a3f467/docs/zh/connectors/sink/HugeGraph.md)
+- [HugeGraph Source 
dev(本文核对版本)](https://github.com/apache/seatunnel/blob/f1a1a0abbe24bdac8cf23307995a78f778a3f467/docs/zh/connectors/source/HugeGraph.md)
+- [JDBC Source](https://seatunnel.apache.org/docs/connectors/source/Jdbc/)
+- [Kafka Source](https://seatunnel.apache.org/docs/connectors/source/Kafka/)
+- [MySQL CDC 
Source](https://seatunnel.apache.org/docs/connectors/source/MySQL-CDC/)
+- [HugeGraph-Loader](/cn/docs/quickstart/toolchain/hugegraph-loader)
+- [HugeGraph-Tools](/cn/docs/quickstart/toolchain/hugegraph-tools)
 - [Apache SeaTunnel GitHub](https://github.com/apache/seatunnel)
-- [HugeGraph GitHub](https://github.com/apache/hugegraph)
+- [Apache HugeGraph GitHub](https://github.com/apache/hugegraph)

Reply via email to