This is an automated email from the ASF dual-hosted git repository.
github-merge-queue[bot] pushed a commit to branch dev
in repository https://gitbox.apache.org/repos/asf/seatunnel.git
The following commit(s) were added to refs/heads/dev by this push:
new eaa6403658 [Docs] Fix option inconsistencies and broken anchors in
connector docs (#12331)
eaa6403658 is described below
commit eaa6403658a0c9c1d7d5b5b5f16a24bfc09974eb
Author: Jast <[email protected]>
AuthorDate: Wed Sep 16 15:09:28 2026 +0000
[Docs] Fix option inconsistencies and broken anchors in connector docs
(#12331)
---
docs/en/connectors/connector-faq.md | 2 +-
docs/en/connectors/sink/BosFile.md | 2 +-
docs/en/connectors/sink/CosFile.md | 2 +-
docs/en/connectors/sink/FtpFile.md | 2 +-
docs/en/connectors/sink/HdfsFile.md | 2 +-
docs/en/connectors/sink/HugeGraph.md | 4 ++--
docs/en/connectors/sink/IoTDBv2.md | 6 +++---
docs/en/connectors/sink/LocalFile.md | 2 +-
docs/en/connectors/sink/OssFile.md | 2 +-
docs/en/connectors/sink/OssJindoFile.md | 2 +-
docs/en/connectors/sink/S3File.md | 2 +-
docs/en/connectors/sink/SftpFile.md | 2 +-
docs/en/connectors/source/Kafka.md | 1 -
docs/en/connectors/source/ObsFile.md | 14 +++++++-------
docs/en/connectors/source/OssFile.md | 2 +-
docs/en/connectors/source/Persistiq.md | 4 ++--
docs/en/edge-agent/configuration.md | 6 +++---
docs/en/edge-agent/faq.md | 4 ++--
docs/en/edge-agent/input-configuration.md | 2 +-
docs/en/edge-agent/output-configuration.md | 2 +-
docs/en/edge-agent/quick-start.md | 8 ++++----
.../configuration/config-encryption-decryption.md | 2 +-
docs/zh/connectors/connector-faq.md | 20 ++++++++++----------
docs/zh/connectors/sink/BosFile.md | 4 ++--
docs/zh/connectors/sink/HugeGraph.md | 4 ++--
docs/zh/connectors/sink/IoTDBv2.md | 6 +++---
docs/zh/connectors/sink/OssJindoFile.md | 2 +-
docs/zh/connectors/source/BosFile.md | 2 +-
docs/zh/connectors/source/FakeSource.md | 2 +-
docs/zh/connectors/source/GcsFile.md | 2 +-
docs/zh/connectors/source/Hbase.md | 2 +-
docs/zh/connectors/source/Kafka.md | 1 -
docs/zh/connectors/source/ObsFile.md | 2 +-
docs/zh/connectors/source/OssFile.md | 2 +-
docs/zh/connectors/source/OssJindoFile.md | 2 +-
docs/zh/connectors/source/Persistiq.md | 6 +++---
docs/zh/connectors/source/Typesense.md | 2 +-
docs/zh/edge-agent/configuration.md | 6 +++---
docs/zh/edge-agent/faq.md | 4 ++--
docs/zh/edge-agent/input-configuration.md | 2 +-
docs/zh/edge-agent/output-configuration.md | 2 +-
docs/zh/edge-agent/quick-start.md | 8 ++++----
docs/zh/engines/zeta/log-analysis-with-ai.md | 2 +-
43 files changed, 78 insertions(+), 80 deletions(-)
diff --git a/docs/en/connectors/connector-faq.md
b/docs/en/connectors/connector-faq.md
index 8fe14a028f..d37b000808 100644
--- a/docs/en/connectors/connector-faq.md
+++ b/docs/en/connectors/connector-faq.md
@@ -49,7 +49,7 @@ Change Data Capture connectors read real-time change events
(INSERT / UPDATE / D
| Connector | Common FAQ Topics |
|---|---|
-| [JDBC Sink](./sink/Jdbc.md#faq) | Automatic table creation, exactly-once
with XA transactions, upsert / primary key configuration, multi-table writing,
missing JDBC driver |
+| [JDBC Sink](./sink/Jdbc.md#troubleshooting) | Automatic table creation,
exactly-once with XA transactions, upsert / primary key configuration,
multi-table writing, missing JDBC driver |
### Data Lakes / File Systems
diff --git a/docs/en/connectors/sink/BosFile.md
b/docs/en/connectors/sink/BosFile.md
index d88f61b950..0e930c479b 100644
--- a/docs/en/connectors/sink/BosFile.md
+++ b/docs/en/connectors/sink/BosFile.md
@@ -66,7 +66,7 @@ To use this connector you need to put `bos-hdfs-sdk` (>=
1.0.4-community) into `
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true.
|
| file_format_type | string | no | "csv"
| File format type, supported: `text`, `csv`,
`parquet`, `orc`, `json`, `excel`, `xml`, `binary`, `canal_json`,
`debezium_json`, `maxwell_json`. |
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv.
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv.
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`.
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used when have_partition is true.
|
diff --git a/docs/en/connectors/sink/CosFile.md
b/docs/en/connectors/sink/CosFile.md
index 52dc1b40f3..7861190e04 100644
--- a/docs/en/connectors/sink/CosFile.md
+++ b/docs/en/connectors/sink/CosFile.md
@@ -66,7 +66,7 @@ To use this connector you need to put
`hadoop-cos-{hadoop.version}-{version}.jar
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true.
|
| file_format_type | string | no | "csv"
| File format type, supported: `text`, `csv`,
`parquet`, `orc`, `json`, `excel`, `xml`, `binary`, `canal_json`,
`debezium_json`, `maxwell_json`. |
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv.
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv.
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`.
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used when have_partition is true.
|
diff --git a/docs/en/connectors/sink/FtpFile.md
b/docs/en/connectors/sink/FtpFile.md
index 3ff9da2fc8..bfa5f671cc 100644
--- a/docs/en/connectors/sink/FtpFile.md
+++ b/docs/en/connectors/sink/FtpFile.md
@@ -55,7 +55,7 @@ By default, we use 2PC commit to ensure `exactly-once`
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true
|
| file_format_type | string | no | "csv"
|
|
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used then have_partition is true
|
diff --git a/docs/en/connectors/sink/HdfsFile.md
b/docs/en/connectors/sink/HdfsFile.md
index 3674741e8d..21d55338d2 100644
--- a/docs/en/connectors/sink/HdfsFile.md
+++ b/docs/en/connectors/sink/HdfsFile.md
@@ -60,7 +60,7 @@ Output data to hdfs file
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when `custom_filename` is `true`.When
the format in the `file_name_expression` parameter is `xxxx-${now}` ,
`filename_time_format` can specify the time format of the path, and the default
value is `yyyy.MM.dd` . The commonly used time formats are listed as
follows:[y:Year,M:Month,d:Day of month,H:Hour in day (0-23),m:Minute in
hour,s:Second in minute] [...]
| file_format_type | string | no | "csv"
| We supported as the following file types:`text`
`csv` `parquet` `orc` `json` `excel` `xml` `binary`.Please note that, The final
file name will end with the file_format's suffix, the suffix of the text file
is `txt`.
[...]
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`
[...]
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format is text and csv,The
separator between columns in a row of data. Only needed by `text` file format.
[...]
+| field_delimiter | string | no | '\001'
| Only used when file_format is text and csv,The
separator between columns in a row of data. Only needed by `text` file format.
[...]
| row_delimiter | string | no | "\n"
| Only used when file_format is text,The separator
between rows in a file. Only needed by `text`, `csv` and `json` file format.
[...]
| have_partition | boolean | no | false
| Whether you need processing partitions.
[...]
| partition_by | array | no | -
| Only used then have_partition is true,Partition
data based on selected fields.
[...]
diff --git a/docs/en/connectors/sink/HugeGraph.md
b/docs/en/connectors/sink/HugeGraph.md
index 9d2e4ce03d..aa1926825c 100644
--- a/docs/en/connectors/sink/HugeGraph.md
+++ b/docs/en/connectors/sink/HugeGraph.md
@@ -45,8 +45,8 @@ New `mappings` configurations default to `schema_save_mode =
CREATE_SCHEMA_WHEN_
| `password` | String | No | - | The password for
HugeGraph authentication. |
| `batch_size` | Integer | No | 500 | The number of
records to buffer before writing to HugeGraph in a single batch. |
| `batch_interval_ms` | Integer | No | 5000 | Retained for
compatibility. To schedule timer flush on Zeta, configure `sink.flush.interval`
in the job `env` block. |
-| `batch_failure_fallback` | Boolean | No | true | When a batch
insert fails, fall back to inserting the batch record by record so a single bad
("poison") record no longer fails the whole batch. Failed records are logged
and skipped; the rest succeed. If every record fails (systemic error), it is
surfaced. Set to `false` to fail the whole batch instead. |
-| `max_insert_errors` | Integer | No | 500 | Maximum number of
records that may be skipped by the single-record fallback
(`batch_failure_fallback=true`) before the task is failed. Bounds the otherwise
unlimited silent skipping of poison records. Set to `-1` for unlimited. Only
applies when `batch_failure_fallback` is enabled. |
+| `batch_failure_fallback` | Boolean | No | false | When a batch
insert fails, fall back to inserting the batch record by record so a single bad
("poison") record no longer fails the whole batch. Failed records are logged
and skipped; the rest succeed. If every record fails (systemic error), it is
surfaced. Set to `false` to fail the whole batch instead. |
+| `max_insert_errors` | Integer | No | 0 | Maximum number of
records that may be skipped by the single-record fallback
(`batch_failure_fallback=true`) before the task is failed. Default `0`: any
skipped record fails the task. Set to `-1` for unlimited. Only applies when
`batch_failure_fallback` is enabled. |
| `failure_data_path` | String | No | - | Optional local
directory. When set, every record skipped by the single-record fallback is
appended (mapped id, label, properties and the server error) to a per-subtask
file (`hugegraph-sink-failures-subtask-N.log`) for offline investigation. In
cluster mode the file is created on the worker node running the sink subtask. |
| `check_vertex` | Boolean | No | false | Whether the
server verifies that an edge's source/target vertices exist when writing edges.
When `false`, edges whose endpoints were never loaded are written as orphan
edges (or trigger server-side phantom vertex auto-creation). Enable to reject
such edges. |
| `max_retries` | Integer | No | 3 | Retries after the
initial attempt. Set to `0` to disable retries. |
diff --git a/docs/en/connectors/sink/IoTDBv2.md
b/docs/en/connectors/sink/IoTDBv2.md
index 0fc2ecc3bd..5fc1770373 100644
--- a/docs/en/connectors/sink/IoTDBv2.md
+++ b/docs/en/connectors/sink/IoTDBv2.md
@@ -57,9 +57,9 @@ Used to write data to IoTDB 2.x. The connector name in job
configuration is `IoT
| key_tag_fields | Array | No | - |
IoTDB-tree: invalid <br/> IoTDB-table: Specify the field names in SeaTunnelRow
to be used as TAG columns
|
| key_attribute_fields | Array | No | - |
IoTDB-tree: invalid <br/> IoTDB-table: Specify the field names in SeaTunnelRow
to be used as ATTRIBUTE columns
|
| batch_size | Integer | No | 1024 |
The connector flushes buffered records to IoTDB when the number of buffered
records reaches `batch_size`. It also flushes before checkpoint commit and when
the writer is closed.
|
-| max_retries | Integer | No | 0 |
The maximum number of retries when flushing records fails.
|
-| retry_backoff_multiplier_ms | Integer | No | 0 |
The multiplier used to calculate the retry backoff delay.
|
-| max_retry_backoff_ms | Integer | No | 0 |
The maximum retry backoff delay in milliseconds.
|
+| max_retries | Integer | No | - |
The maximum number of retries when flushing records fails.
|
+| retry_backoff_multiplier_ms | Integer | No | - |
The multiplier used to calculate the retry backoff delay.
|
+| max_retry_backoff_ms | Integer | No | - |
The maximum retry backoff delay in milliseconds.
|
| default_thrift_buffer_size | Integer | No | - |
Thrift init buffer size in IoTDB client
|
| max_thrift_frame_size | Integer | No | - |
Thrift max frame size in IoTDB client
|
| zone_id | String | No | - |
java.time.ZoneId in IoTDB client
|
diff --git a/docs/en/connectors/sink/LocalFile.md
b/docs/en/connectors/sink/LocalFile.md
index 61ec2282f9..c1758da1b7 100644
--- a/docs/en/connectors/sink/LocalFile.md
+++ b/docs/en/connectors/sink/LocalFile.md
@@ -59,7 +59,7 @@ By default, we use 2PC commit to ensure `exactly-once`
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true.
|
| file_format_type | string | no | "csv"
| File format type, supported: `text`, `csv`,
`parquet`, `orc`, `json`, `excel`, `xml`, `binary`, `canal_json`,
`debezium_json`, `maxwell_json`. |
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv.
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv.
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`.
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used when have_partition is true.
|
diff --git a/docs/en/connectors/sink/OssFile.md
b/docs/en/connectors/sink/OssFile.md
index 81688c68df..c84668eb16 100644
--- a/docs/en/connectors/sink/OssFile.md
+++ b/docs/en/connectors/sink/OssFile.md
@@ -108,7 +108,7 @@ If write to `csv`, `text`, `json` file type, All column
will be string.
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true.
|
| file_format_type | string | no | "csv"
| File format type, supported: `text`, `csv`,
`parquet`, `orc`, `json`, `excel`, `xml`, `binary`, `canal_json`,
`debezium_json`, `maxwell_json`. |
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`.
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv.
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv.
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`.
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used when have_partition is true.
|
diff --git a/docs/en/connectors/sink/OssJindoFile.md
b/docs/en/connectors/sink/OssJindoFile.md
index db08db290d..f9622a54b3 100644
--- a/docs/en/connectors/sink/OssJindoFile.md
+++ b/docs/en/connectors/sink/OssJindoFile.md
@@ -74,7 +74,7 @@ The connector targets Alibaba OSS via the Jindo SDK. The
Jindo SDK jars (`jindo-
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true.
|
| file_format_type | string | no | "csv"
| File format type, supported: `text`, `csv`,
`parquet`, `orc`, `json`, `excel`, `xml`, `binary`, `canal_json`,
`debezium_json`, `maxwell_json`. |
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`.
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv.
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv.
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`.
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used when have_partition is true.
|
diff --git a/docs/en/connectors/sink/S3File.md
b/docs/en/connectors/sink/S3File.md
index 3ec40bf246..84e073bf7d 100644
--- a/docs/en/connectors/sink/S3File.md
+++ b/docs/en/connectors/sink/S3File.md
@@ -118,7 +118,7 @@ If write to `csv`, `text` file type, All column will be
string.
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true
|
| file_format_type | string | no | "csv"
|
|
| filename_extension | string | no | -
| Override the default file name
extensions with custom file name extensions. E.g. `.xml`, `.json`, `dat`,
`.customtype` |
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format is text and
csv
|
+| field_delimiter | string | no | '\001'
| Only used when file_format is text and csv
|
| row_delimiter | string | no | "\n"
| Only used when file_format is `text`,
`csv` and `json`
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used when have_partition is true
|
diff --git a/docs/en/connectors/sink/SftpFile.md
b/docs/en/connectors/sink/SftpFile.md
index 25378cf860..7ba90dca1f 100644
--- a/docs/en/connectors/sink/SftpFile.md
+++ b/docs/en/connectors/sink/SftpFile.md
@@ -59,7 +59,7 @@ If you use SeaTunnel Engine, It automatically integrated the
hadoop jar when you
| filename_time_format | string | no | "yyyy.MM.dd"
| Only used when custom_filename is true
|
| file_format_type | string | no | "csv"
|
|
| filename_extension | string | no | -
| Override the default file name extensions with
custom file name extensions. E.g. `.xml`, `.json`, `dat`, `.customtype`
|
-| field_delimiter | string | no | '\001' for text
and ',' for csv | Only used when file_format_type is text and csv
|
+| field_delimiter | string | no | '\001'
| Only used when file_format_type is text and csv
|
| row_delimiter | string | no | "\n"
| Only used when file_format_type is `text`, `csv`
and `json`
|
| have_partition | boolean | no | false
| Whether you need processing partitions.
|
| partition_by | array | no | -
| Only used then have_partition is true
|
diff --git a/docs/en/connectors/source/Kafka.md
b/docs/en/connectors/source/Kafka.md
index b282bf42f5..4cfb662557 100644
--- a/docs/en/connectors/source/Kafka.md
+++ b/docs/en/connectors/source/Kafka.md
@@ -84,7 +84,6 @@ They can be downloaded via install-plugin.sh or from the
Maven central repositor
| protobuf_schema | String
| No | - |
Effective when the format is set to protobuf, specifies the Schema definition
[...]
| strip_schema_registry_header | Boolean
| No | false |
Effective when the format is set to protobuf or avro. For protobuf, strips the
Confluent Schema Registry header before deserialization. For avro, strips the
fixed five-byte header (magic byte and schema ID); `avro_schema` is required
when enabled, and no Schema Registry lookup is performed. |
| reader_cache_queue_size | Integer
| No | 2 |
The capacity of the fetcher-to-reader element queue. Each element is one
`consumer.poll()` batch, not a single message. See
[reader_cache_queue_size](#reader_cache_queue_size) for details. |
-| is_native | Boolean
| No | false |
Supports retaining the source information of the record.
[...]
| kafka_headers_fields | Array
| No | - |
Specify which Kafka message header keys to extract as row fields. Each header
value is read as a STRING type and appended to the output row after the regular
schema fields. Cannot be used with NATIVE format.
[...]
> On restore from checkpoint or savepoint, Kafka Source resumes from the
> checkpointed split offsets.
diff --git a/docs/en/connectors/source/ObsFile.md
b/docs/en/connectors/source/ObsFile.md
index 7fd14e4d45..496bd9b869 100644
--- a/docs/en/connectors/source/ObsFile.md
+++ b/docs/en/connectors/source/ObsFile.md
@@ -72,7 +72,7 @@ It only supports hadoop version **2.9.X+**.
| access_secret | string | yes | - | The
access secret of obs file system
|
| endpoint | string | yes | - | The
endpoint of obs file system
|
| read_columns | list | no | - | The
read column list of the data source, user can use it to implement field
projection.[Tips](#read_columns)
|
-| delimiter | string | no | \001 |
Field delimiter, used to tell connector how to slice and dice fields when
reading text files
|
+| delimiter/field_delimiter | string | no | \001 |
Field delimiter, used to tell connector how to slice and dice fields when
reading text files. Default `\001`, the same as hive's default delimiter.
**delimiter** parameter will deprecate after version 2.3.5, please use
**field_delimiter** instead.
|
| row_delimiter | string | no | \n | Row
delimiter, used to tell connector how to slice and dice rows when reading text
files. Default is `\n` for text files.
|
| parse_partition_from_path | boolean | no | true |
Control whether parse the partition keys and values from file path.
[Tips](#parse_partition_from_path)
|
| skip_header_row_number | long | no | 0 | Skip
the first few lines, but only for the txt and csv.
|
@@ -203,13 +203,13 @@ tyrantlucifer#26#male
|-----------------------|
| tyrantlucifer#26#male |
-> If you assign data schema, you should also assign the option `delimiter` too
except CSV file type
+> If you assign data schema, you should also assign the option
`field_delimiter` too except CSV file type
>
-> you should assign schema and delimiter as the following:
+> you should assign schema and field_delimiter as the following:
```hocon
-delimiter = "#"
+field_delimiter = "#"
schema {
fields {
name = string
@@ -282,7 +282,7 @@ schema {
> The schema of upstream data. For more details, please refer to [Schema
> Feature](../../introduction/concepts/schema-feature.md).
-#### <span id="schema"> read_columns </span>
+#### <span id="read_columns"> read_columns </span>
> The read column list of the data source, user can use it to implement field
> projection.
>
@@ -297,7 +297,7 @@ schema {
> If the user wants to use this feature when reading `text` `json` `csv`
> files, the schema option must be configured
-#### <span id="common_options "> common options </span>
+#### <span id="common_options"> common options </span>
> Source plugin common parameters, please refer to [Source Common
> Options](../common-options/source-common-options.md) for details.
@@ -409,7 +409,7 @@ schema {
access_secret = "xxxxxxxxxxxxxxxxxxxxxx"
endpoint = "obs.xxxxxx.myhuaweicloud.com"
file_format_type = "csv"
- delimiter = ","
+ field_delimiter = ","
}
```
diff --git a/docs/en/connectors/source/OssFile.md
b/docs/en/connectors/source/OssFile.md
index 15df807928..6db8bf754d 100644
--- a/docs/en/connectors/source/OssFile.md
+++ b/docs/en/connectors/source/OssFile.md
@@ -193,7 +193,7 @@ If you assign file type to `parquet` `orc`, schema option
not required, connecto
| read_columns | list | no | - | The
read column list of the data source, user can use it to implement field
projection. The file type supported column projection as the following shown:
`text` `csv` `parquet` `orc` `json` `excel` `xml` . If the user wants to use
this feature when reading `text` `json` `csv` files, the "schema" option must
be configured. |
| access_key | string | no | - |
|
| access_secret | string | no | - |
|
-| delimiter | string | no | \001 |
Field delimiter, used to tell connector how to slice and dice fields when
reading text files. Default `\001`, the same as hive's default delimiter.
|
+| delimiter/field_delimiter | string | no | \001 |
Field delimiter, used to tell connector how to slice and dice fields when
reading text files. Default `\001`, the same as hive's default delimiter.
**delimiter** parameter will deprecate after version 2.3.5, please use
**field_delimiter** instead.
|
| row_delimiter | string | no | \n | Row
delimiter, used to tell connector how to slice and dice rows when reading text
files. Default `\n`.
|
| parse_partition_from_path | boolean | no | true |
Control whether parse the partition keys and values from file path. For example
if you read a file from path
`oss://hadoop-cluster/tmp/seatunnel/parquet/name=tyrantlucifer/age=26`. Every
record data from file will be added these two fields: name="tyrantlucifer",
age=16 |
| date_format | string | no | yyyy-MM-dd | Date
type format, used to tell connector how to convert string to date, supported as
the following formats:`yyyy-MM-dd` `yyyy.MM.dd` `yyyy/MM/dd`. default
`yyyy-MM-dd`
|
diff --git a/docs/en/connectors/source/Persistiq.md
b/docs/en/connectors/source/Persistiq.md
index 9b71372ae4..ebc16545da 100644
--- a/docs/en/connectors/source/Persistiq.md
+++ b/docs/en/connectors/source/Persistiq.md
@@ -30,7 +30,7 @@ Used to read data from Persistiq.
| params | Map | No | - |
| body | String | No | - |
| json_field | Config | No | - |
-| content_json | String | No | - |
+| content_field | String | No | - |
| poll_interval_millis | int | No | - |
| retry | int | No | - |
| retry_backoff_multiplier_ms | int | No | 100 |
@@ -134,7 +134,7 @@ connector will generate data as the following:
The schema fields of upstream data. For more details, please refer to [Schema
Feature](../../introduction/concepts/schema-feature.md).
-### content_json [String]
+### content_field [String]
This parameter can get some json data.If you only need the data in the 'book'
section, configure `content_field = "$.store.book.*"`.
diff --git a/docs/en/edge-agent/configuration.md
b/docs/en/edge-agent/configuration.md
index aafcfa9255..4de35cea5e 100644
--- a/docs/en/edge-agent/configuration.md
+++ b/docs/en/edge-agent/configuration.md
@@ -34,7 +34,7 @@ Process-wide settings.
| Key | Type | Required | Default | Description
|
| -------------------- | ------ | -------- | ------------- |
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|
-| id | string | No | auto | Agent instance
identity (logs, ops). Auto: see [Identity file](#identity-file-edge-agentid).
|
+| id | string | No | auto | Agent instance
identity (logs, ops). Auto: see [Identity file](#identity-file).
|
| delivery-guarantee | string | No | BEST_EFFORT | Outbound delivery
mode. BEST_EFFORT (aliases: best-effort, best_effort): durable local WAL with
retry; the same record may reach output more than once — design downstream
consumers to be idempotent. NON (aliases: non, none): no WAL, no persistence;
events are sent directly from memory and dropped on failure; completely
stateless, no local persistence dependency. Details: [Delivery
mode](./architecture-overview.md#62-delivery-mode). |
| idle-sleep-ms | long | No | 200 | Sleep (ms) when a
scheduler loop iteration makes no progress. Must be > 0. |
| bulk-max-size | integer | No | 256 | Flush in-memory
reader buffer to WAL when this many events are pending. |
@@ -67,7 +67,7 @@ For glob patterns, multiline log assembly, and scenario-based
YAML examples, see
| Key | Type | Required | Default | Description
|
| ----------------------- | -------------- | -------- | -------- |
---------------------------------------------------------------------------------------------------------------------------------------------
|
-| id | string | No | auto | Source
identity; WAL/position sourceId. Auto: see [Identity
file](#identity-file-edge-agentid).
|
+| id | string | No | auto | Source
identity; WAL/position sourceId. Auto: see [Identity file](#identity-file).
|
| type | string | No | file | Input plugin
id. Only file is implemented.
|
| paths | list of string | Yes | — | Glob patterns for
files to tail (e.g. /var/log/*.log). Must be non-empty; blank entries are
invalid. |
| encoding | string | No | UTF-8 | File character
encoding.
|
@@ -164,7 +164,7 @@ For endpoint/token alignment, RAW vs PACKET, and scenario
YAML, see [Output Conf
| Key | Type | Required | Default | Description
|
| ------ | ------ | -------- | --------- |
-----------------------------------------------------------------------------------------------------------------------
|
-| id | string | No | auto | Outbound logical identity (logs,
migration; not on the wire). Auto: see [Identity
file](#identity-file-edge-agentid). |
+| id | string | No | auto | Outbound logical identity (logs,
migration; not on the wire). Auto: see [Identity file](#identity-file). |
| type | string | No | console | transport: EdgeSocket client. console:
write payload logs as EDGE_CONSOLE_OUTPUT to log/edge-agent.log (debug). |
diff --git a/docs/en/edge-agent/faq.md b/docs/en/edge-agent/faq.md
index 5eec7a0934..2d0a6c4f85 100644
--- a/docs/en/edge-agent/faq.md
+++ b/docs/en/edge-agent/faq.md
@@ -11,7 +11,7 @@ High-frequency questions about identity, migration, WAL DEAD
rows, and multiple
### What is edge-agent.id?
-An identity file at the install root (next to edge-agent.pid). When agent.id,
input.id, or output.id are omitted from YAML, the agent reads or writes them
here; explicit YAML IDs win. See [Identity
file](configuration.md#identity-file-edge-agentid).
+An identity file at the install root (next to edge-agent.pid). When agent.id,
input.id, or output.id are omitted from YAML, the agent reads or writes them
here; explicit YAML IDs win. See [Identity
file](configuration.md#identity-file).
### What do agent.id, input.id, and output.id do?
@@ -46,7 +46,7 @@ Launcher environment variable; default
$EDGE_AGENT_HOME/edge-agent.id. See [Oper
## WAL persistence
-Default path, on-disk files (data, -wal, -shm), and what each table stores:
[Configuration — WAL persistence
files](configuration.md#sqlite-persistence-files). Console mode still uses this
database.
+Default path, on-disk files (data, -wal, -shm), and what each table stores:
[Configuration — WAL persistence
files](configuration.md#wal-persistence-files). Console mode still uses this
database.
### What does WAL status DEAD mean?
diff --git a/docs/en/edge-agent/input-configuration.md
b/docs/en/edge-agent/input-configuration.md
index 9bfa869584..53cb9a3d2e 100644
--- a/docs/en/edge-agent/input-configuration.md
+++ b/docs/en/edge-agent/input-configuration.md
@@ -211,7 +211,7 @@ input:
### 8. Stable input.id across reinstalls
-Omitting id uses a stable source identity as the WAL / position sourceId (see
[Identity file](configuration.md#identity-file-edge-agentid)). Set an explicit
id for multiple agents or WAL restore; keep edge-agent.id with the WAL database:
+Omitting id uses a stable source identity as the WAL / position sourceId (see
[Identity file](configuration.md#identity-file)). Set an explicit id for
multiple agents or WAL restore; keep edge-agent.id with the WAL database:
```yaml
input:
diff --git a/docs/en/edge-agent/output-configuration.md
b/docs/en/edge-agent/output-configuration.md
index 1a3b5d3c96..ffc72ffa51 100644
--- a/docs/en/edge-agent/output-configuration.md
+++ b/docs/en/edge-agent/output-configuration.md
@@ -175,4 +175,4 @@ Transport reconnect is separate from WAL retry.* (scheduler
/ outbound row repla
## output.id and migration
-output.id labels the outbound side for logs and migration only (not on the
wire). See [Identity file](configuration.md#identity-file-edge-agentid). When
moving hosts, copy edge-agent.id with the WAL database and config/agent.yaml.
+output.id labels the outbound side for logs and migration only (not on the
wire). See [Identity file](configuration.md#identity-file). When moving hosts,
copy edge-agent.id with the WAL database and config/agent.yaml.
diff --git a/docs/en/edge-agent/quick-start.md
b/docs/en/edge-agent/quick-start.md
index 56415adc6d..5379ac2e7a 100644
--- a/docs/en/edge-agent/quick-start.md
+++ b/docs/en/edge-agent/quick-start.md
@@ -10,8 +10,8 @@ Two validation paths:
| Path | Engine
required | Purpose |
| --------------------------------------------------------------- |
--------------- | ------------------------------------------ |
-| [Console local test](#console-local-test-no-engine) | No
| Install, collect, local WAL without Engine |
-| [Production setup (with Engine)](#production-setup-with-engine) | Yes
| Agent → Engine job |
+| [Console local test](#console-local-test) | No |
Install, collect, local WAL without Engine |
+| [Production setup (with Engine)](#production-setup) | Yes |
Agent → Engine job |
For production hardening, see [Deployment Guide](deployment-guide.md) and
[Operations](operations.md).
@@ -79,7 +79,7 @@ Search log/edge-agent.log for EDGE_CONSOLE_OUTPUT (console
writes via the app lo
:::
-The agent still creates edge-agent.id and the WAL persistence files under
data/ (wal.db, wal.db-wal, wal.db-shm), even without Engine. See [Configuration
— WAL persistence files](configuration.md#sqlite-persistence-files).
+The agent still creates edge-agent.id and the WAL persistence files under
data/ (wal.db, wal.db-wal, wal.db-shm), even without Engine. See [Configuration
— WAL persistence files](configuration.md#wal-persistence-files).
### Stop
@@ -87,7 +87,7 @@ The agent still creates edge-agent.id and the WAL persistence
files under data/
sh bin/seatunnel-edge-agent.sh stop
```
-To continue with Engine, see [Production setup](#production-setup-with-engine).
+To continue with Engine, see [Production setup](#production-setup).
## Production setup
diff --git a/docs/en/introduction/configuration/config-encryption-decryption.md
b/docs/en/introduction/configuration/config-encryption-decryption.md
index 64327bae38..432b114bea 100644
--- a/docs/en/introduction/configuration/config-encryption-decryption.md
+++ b/docs/en/introduction/configuration/config-encryption-decryption.md
@@ -6,7 +6,7 @@ In most production environments, sensitive configuration items
such as passwords
## How to use
-SeaTunnel comes with the function of base64 encryption and decryption, but it
is not recommended for production use, it is recommended that users implement
custom encryption and decryption logic. You can refer to this chapter [How to
implement user-defined encryption and decryption](#How to implement
user-defined encryption and decryption) get more details about it.
+SeaTunnel comes with the function of base64 encryption and decryption, but it
is not recommended for production use, it is recommended that users implement
custom encryption and decryption logic. You can refer to this chapter [How to
implement user-defined encryption and
decryption](#how-to-implement-user-defined-encryption-and-decryption) get more
details about it.
Base64 encryption support encrypt the following parameters by default:
- username
diff --git a/docs/zh/connectors/connector-faq.md
b/docs/zh/connectors/connector-faq.md
index 3a0de3eef0..e812bdda3f 100644
--- a/docs/zh/connectors/connector-faq.md
+++ b/docs/zh/connectors/connector-faq.md
@@ -18,9 +18,9 @@ CDC(Change Data Capture)连接器从数据库事务日志中读取实时变
| 连接器 | 常见 FAQ 主题 |
|---|---|
-| [MySQL CDC](./source/MySQL-CDC.md#faq) | 所需权限、binlog 配置、从库支持、无主键表、快照阶段、DDL
传播、`server-id` 冲突、快照性能、时区/字符集 |
-| [PostgreSQL CDC](./source/PostgreSQL-CDC.md#faq) |
所需权限、逻辑解码插件、从库支持、无主键表、复制槽管理、复制延迟 |
-| [Oracle CDC](./source/Oracle-CDC.md#faq) | LogMiner 权限、补充日志、CDB/PDB
多租户、无主键表、LogMiner 性能、支持的 Oracle 版本 |
+| [MySQL CDC](./source/MySQL-CDC.md#常见问题) | 所需权限、binlog 配置、从库支持、无主键表、快照阶段、DDL
传播、`server-id` 冲突、快照性能、时区/字符集 |
+| [PostgreSQL CDC](./source/PostgreSQL-CDC.md#常见问题) |
所需权限、逻辑解码插件、从库支持、无主键表、复制槽管理、复制延迟 |
+| [Oracle CDC](./source/Oracle-CDC.md#常见问题) | LogMiner 权限、补充日志、CDB/PDB
多租户、无主键表、LogMiner 性能、支持的 Oracle 版本 |
---
@@ -28,8 +28,8 @@ CDC(Change Data Capture)连接器从数据库事务日志中读取实时变
| 连接器 | 常见 FAQ 主题 |
|---|---|
-| [Kafka Source](./source/Kafka.md#faq) | `start_mode` 选项、按消息 key
过滤、支持的格式、SASL/Kerberos 认证、消费组 offset 提交 |
-| [Kafka Sink](./sink/Kafka.md#faq) | 自动创建 Topic、`partition_key_fields`
行为、精确一次投递、SASL/Kerberos 认证、支持的格式 |
+| [Kafka Source](./source/Kafka.md#常见问题) | `start_mode` 选项、按消息 key
过滤、支持的格式、SASL/Kerberos 认证、消费组 offset 提交 |
+| [Kafka Sink](./sink/Kafka.md#常见问题) | 自动创建 Topic、`partition_key_fields`
行为、精确一次投递、SASL/Kerberos 认证、支持的格式 |
---
@@ -39,21 +39,21 @@ CDC(Change Data Capture)连接器从数据库事务日志中读取实时变
| 连接器 | 常见 FAQ 主题 |
|---|---|
-| [Doris Sink](./sink/Doris.md#faq) | 自动建表、2PC 精确一次、"Label already exists"
错误、DELETE 传播、列名大小写、Stream Load 格式 |
-| [StarRocks Sink](./sink/StarRocks.md#faq) | 自动建表、Upsert 与 DELETE
支持、`labelPrefix` 用法、列名大小写、`nodeUrls` 与 `base-url` |
-| [ClickHouse Sink](./sink/Clickhouse.md#faq) | 自动建表、批量写入性能、支持的数据类型、"Table
doesn't exist" 错误 |
+| [Doris Sink](./sink/Doris.md#常见问题) | 自动建表、2PC 精确一次、"Label already exists"
错误、DELETE 传播、列名大小写、Stream Load 格式 |
+| [StarRocks Sink](./sink/StarRocks.md#常见问题) | 自动建表、Upsert 与 DELETE
支持、`labelPrefix` 用法、列名大小写、`nodeUrls` 与 `base-url` |
+| [ClickHouse Sink](./sink/Clickhouse.md#常见问题) | 自动建表、批量写入性能、支持的数据类型、"Table
doesn't exist" 错误 |
### 关系型数据库
| 连接器 | 常见 FAQ 主题 |
|---|---|
-| [JDBC Sink](./sink/Jdbc.md#faq) | 自动建表、XA 事务精确一次、Upsert / 主键配置、多表写入、缺少 JDBC
驱动 |
+| [JDBC Sink](./sink/Jdbc.md#故障排查) | 自动建表、XA 事务精确一次、Upsert / 主键配置、多表写入、缺少 JDBC
驱动 |
### 数据湖 / 文件系统
| 连接器 | 常见 FAQ 主题 |
|---|---|
-| [Hive Sink](./sink/Hive.md#faq) | 支持的文件格式、分区表、Kerberos 认证、小文件问题、Schema 演进 |
+| [Hive Sink](./sink/Hive.md#常见问题) | 支持的文件格式、分区表、Kerberos 认证、小文件问题、Schema 演进 |
---
diff --git a/docs/zh/connectors/sink/BosFile.md
b/docs/zh/connectors/sink/BosFile.md
index f904891984..9776b11596 100644
--- a/docs/zh/connectors/sink/BosFile.md
+++ b/docs/zh/connectors/sink/BosFile.md
@@ -26,7 +26,7 @@ import ChangeLog from '../changelog/connector-file-bos.md';
## 主要特性
-- [x] [多模态](../../introduction/concepts/connector-v2-features.md#multimodal)
+- [x] [多模态](../../introduction/concepts/connector-v2-features.md#多模态multimodal)
使用 binary 格式读写任意类型文件,可将任意文件同步到目标位置。
@@ -66,7 +66,7 @@ import ChangeLog from '../changelog/connector-file-bos.md';
| filename_time_format | string | 否 | "yyyy.MM.dd"
| custom_filename 为 true 时使用 |
| file_format_type | string | 否 | "csv"
| 支持 text、csv、parquet、orc、json、excel、xml、binary 等 |
| filename_extension | string | 否 | -
| 自定义文件扩展名 |
-| field_delimiter | string | 否 | text 为 \001,csv 为 ,
| text/csv 格式使用 |
+| field_delimiter | string | 否 | '\001'
| text/csv 格式使用 |
| row_delimiter | string | 否 | "\n"
| text/csv/json 格式使用 |
| have_partition | boolean | 否 | false
| 是否按分区写入 |
| partition_by | array | 否 | -
| have_partition 为 true 时使用 |
diff --git a/docs/zh/connectors/sink/HugeGraph.md
b/docs/zh/connectors/sink/HugeGraph.md
index b914a3a202..f8b64a324f 100644
--- a/docs/zh/connectors/sink/HugeGraph.md
+++ b/docs/zh/connectors/sink/HugeGraph.md
@@ -45,8 +45,8 @@ HugeGraph sink连接器允许您将数据从SeaTunnel写入Apache HugeGraph,
| `password` | String | 否 | - | 用于HugeGraph身份验证的密码。
|
| `batch_size` | Integer | 否 | 500 | 在单批次写入HugeGraph之前缓冲的记录数。
|
| `batch_interval_ms` | Integer | 否 | 5000 | 为兼容性保留。在 Zeta
上需要定时刷新时,请在作业 `env` 中配置 `sink.flush.interval`。 |
-| `batch_failure_fallback` | Boolean | 否 | true |
批量写入失败时,降级为逐条写入,使单条“毒药”记录不再拖垮整批。失败记录会记录日志并跳过,其余成功;若整批全部失败(系统性错误)则抛出。设为 `false`
则整批失败。 |
-| `max_insert_errors` | Integer | 否 | 500 |
逐条降级(`batch_failure_fallback=true`)累计跳过的失败记录达到该数量后使任务失败,用于约束原本无上限的“毒药”记录静默跳过。设为
`-1` 表示不限。仅在开启 `batch_failure_fallback` 时生效。 |
+| `batch_failure_fallback` | Boolean | 否 | false |
批量写入失败时,降级为逐条写入,使单条“毒药”记录不再拖垮整批。失败记录会记录日志并跳过,其余成功;若整批全部失败(系统性错误)则抛出。设为 `false`
则整批失败。 |
+| `max_insert_errors` | Integer | 否 | 0 |
逐条降级(`batch_failure_fallback=true`)累计跳过的失败记录达到该数量后使任务失败。默认
`0`:任何被跳过的记录都会使任务失败。设为 `-1` 表示不限。仅在开启 `batch_failure_fallback` 时生效。 |
| `failure_data_path` | String | 否 | - |
可选本地目录。设置后,逐条降级跳过的每条记录(映射后的
id、label、属性及服务端错误)会追加写入按子任务区分的文件(`hugegraph-sink-failures-subtask-N.log`)以便离线排查。集群模式下文件写在运行该
sink 子任务的 worker 节点上。 |
| `check_vertex` | Boolean | 否 | false | 写入边时服务端是否校验边的源/目标顶点是否存在。为
`false` 时,端点从未写入的边会被写成孤儿边(或触发服务端幻影顶点自动创建)。开启后此类边会被拒绝。 |
| `max_retries` | Integer | 否 | 3 | 首次请求失败后的重试次数。设置为 `0`
可禁用重试。 |
diff --git a/docs/zh/connectors/sink/IoTDBv2.md
b/docs/zh/connectors/sink/IoTDBv2.md
index c529c20550..0668715778 100644
--- a/docs/zh/connectors/sink/IoTDBv2.md
+++ b/docs/zh/connectors/sink/IoTDBv2.md
@@ -57,9 +57,9 @@ import ChangeLog from '../changelog/connector-iotdb.md';
| key_tag_fields | Array | 否 | - | IoTDB 树模型:不生效;<br/>
IoTDB 表模型:在 SeaTunnelRow 中指定 IoTDB 标签列(TAG)的字段名
|
| key_attribute_fields | Array | 否 | - | IoTDB 树模型:不生效;<br/>
IoTDB 表模型:在 SeaTunnelRow 中指定 IoTDB 属性列(ATTRIBUTE)的字段名
|
| batch_size | Integer | 否 | 1024 | 缓存的记录数达到
`batch_size` 时,连接器会把数据刷新到 IoTDB;在 checkpoint 提交前和写入器关闭时也会刷新。
|
-| max_retries | Integer | 否 | 0 | 刷新失败时的最大重试次数。
|
-| retry_backoff_multiplier_ms | Integer | 否 | 0 | 计算重试等待时间的退避倍数,单位为毫秒。
|
-| max_retry_backoff_ms | Integer | 否 | 0 | 最大重试等待时间,单位为毫秒。
|
+| max_retries | Integer | 否 | - | 刷新失败时的最大重试次数。
|
+| retry_backoff_multiplier_ms | Integer | 否 | - | 计算重试等待时间的退避倍数,单位为毫秒。
|
+| max_retry_backoff_ms | Integer | 否 | - | 最大重试等待时间,单位为毫秒。
|
| default_thrift_buffer_size | Integer | 否 | - | IoTDB 客户端使用的默认
Thrift 缓冲区大小。
|
| max_thrift_frame_size | Integer | 否 | - | IoTDB 客户端使用的最大
Thrift 帧大小。
|
| zone_id | String | 否 | - | IoTDB 客户端使用的
`java.time.ZoneId`。
|
diff --git a/docs/zh/connectors/sink/OssJindoFile.md
b/docs/zh/connectors/sink/OssJindoFile.md
index 7210531da9..1eb5d434f6 100644
--- a/docs/zh/connectors/sink/OssJindoFile.md
+++ b/docs/zh/connectors/sink/OssJindoFile.md
@@ -74,7 +74,7 @@ import ChangeLog from
'../changelog/connector-file-oss-jindo.md';
| filename_time_format | string | 否 | "yyyy.MM.dd"
| 仅在 `custom_filename` 为 `true` 时使用。
|
| file_format_type | string | 否 | "csv"
|
文件格式类型,支持:`text`、`csv`、`parquet`、`orc`、`json`、`excel`、`xml`、`binary`、`canal_json`、`debezium_json`、`maxwell_json`。
|
| filename_extension | string | 否 | -
| 使用自定义文件扩展名覆盖默认扩展名,例如 `.xml`、`.json`、`dat`、`.customtype`。
|
-| field_delimiter | string | 否 | '\001' for text and
',' for csv | 仅当 `file_format_type` 为 `text` 和 `csv` 时使用。
|
+| field_delimiter | string | 否 | '\001'
| 仅当 `file_format_type` 为 `text` 和 `csv` 时使用。
|
| row_delimiter | string | 否 | "\n"
| 仅当 `file_format_type` 为 `text`、`csv` 和 `json` 时使用。
|
| have_partition | boolean | 否 | false
| 是否需要处理分区。
|
| partition_by | array | 否 | -
| 仅在 `have_partition` 为 `true` 时使用。
|
diff --git a/docs/zh/connectors/source/BosFile.md
b/docs/zh/connectors/source/BosFile.md
index a4f8ee6b0b..5f6bc3cf73 100644
--- a/docs/zh/connectors/source/BosFile.md
+++ b/docs/zh/connectors/source/BosFile.md
@@ -14,7 +14,7 @@ import ChangeLog from '../changelog/connector-file-bos.md';
- [x] [批处理](../../introduction/concepts/connector-v2-features.md)
- [ ] [流处理](../../introduction/concepts/connector-v2-features.md)
-- [x] [多模态](../../introduction/concepts/connector-v2-features.md#multimodal)
+- [x] [多模态](../../introduction/concepts/connector-v2-features.md#多模态multimodal)
使用 binary 格式读写任意类型文件(视频、图片等),可将任意文件同步到目标位置。
diff --git a/docs/zh/connectors/source/FakeSource.md
b/docs/zh/connectors/source/FakeSource.md
index fdc1608df1..cb03571f08 100644
--- a/docs/zh/connectors/source/FakeSource.md
+++ b/docs/zh/connectors/source/FakeSource.md
@@ -77,7 +77,7 @@ FakeSource 是一个虚拟数据源,它根据用户定义的 schema 数据结
### 简单示例
-> 此示例随机生成指定类型的数据。如果您想了解如何声明字段类型,请点击
[这里](../../introduction/concepts/schema-feature.md#how-to-declare-type-supported)。
+> 此示例随机生成指定类型的数据。如果您想了解如何声明字段类型,请点击
[这里](../../introduction/concepts/schema-feature.md#如何声明支持的类型)。
```hocon
schema = {
diff --git a/docs/zh/connectors/source/GcsFile.md
b/docs/zh/connectors/source/GcsFile.md
index dd58c949cd..2578c029cd 100644
--- a/docs/zh/connectors/source/GcsFile.md
+++ b/docs/zh/connectors/source/GcsFile.md
@@ -16,7 +16,7 @@ import ChangeLog from '../changelog/connector-file-gcs.md';
- [x] [精确一次](../../introduction/concepts/connector-v2-features.md)
- [x] [列投影](../../introduction/concepts/connector-v2-features.md)
- [x] [并行度](../../introduction/concepts/connector-v2-features.md)
-- [x] [多模态](../../introduction/concepts/connector-v2-features.md#multimodal)
+- [x] [多模态](../../introduction/concepts/connector-v2-features.md#多模态multimodal)
- [x] 多表 Source
- [x]
文件格式:`text`、`csv`、`parquet`、`orc`、`json`、`excel`、`xml`、`binary`、`markdown` 和
`pdf`
diff --git a/docs/zh/connectors/source/Hbase.md
b/docs/zh/connectors/source/Hbase.md
index 72c9475e15..6a6b897133 100644
--- a/docs/zh/connectors/source/Hbase.md
+++ b/docs/zh/connectors/source/Hbase.md
@@ -58,7 +58,7 @@ HBase 的 zookeeper 集群主机,例如:“hadoop001:2181,hadoop002:2181,had
HBase 使用字节数组进行存储。因此,您需要为表中的每一列配置数据类型。
行键列使用 `rowkey`,HBase 单元格使用 `列簇:列名` 形式,例如 `info:name`。
-更多信息请参阅:[模式声明指南](../../introduction/concepts/schema-feature.md#how-to-declare-type-supported)。
+更多信息请参阅:[模式声明指南](../../introduction/concepts/schema-feature.md#如何声明支持的类型)。
### hbase_extra_config [config]
diff --git a/docs/zh/connectors/source/Kafka.md
b/docs/zh/connectors/source/Kafka.md
index 3d064dbe83..c5f65c1ffc 100644
--- a/docs/zh/connectors/source/Kafka.md
+++ b/docs/zh/connectors/source/Kafka.md
@@ -83,7 +83,6 @@ header 字段和事件时间元数据。
| protobuf_schema | String |
否 | - | 当格式设置为 protobuf 时有效,指定 Schema 定义。
|
| strip_schema_registry_header | Boolean |
否 | false | 当格式设置为 protobuf 或 avro 时有效。protobuf
会在反序列化前去除 Confluent Schema Registry 头;avro 会去除固定的 5 字节头(magic byte 和 schema
ID)。avro 启用此选项时必须同时配置 `avro_schema`,且不会查询 Schema Registry。 |
| reader_cache_queue_size | Integer |
否 | 2 | Fetcher 与 Reader 线程之间缓冲队列的容量。每个元素是一次
`consumer.poll()` 的整批结果,而非单条消息。详见
[reader_cache_queue_size](#reader_cache_queue_size)。 |
-| is_native | Boolean |
否 | false | 支持保留record的源信息。
|
| kafka_headers_fields | Array |
否 | - | 指定要从 Kafka 消息 header 中提取并映射为行字段的 header
key 列表。每个 header 值以 STRING 类型追加到输出行的末尾(位于正常 schema 字段之后)。不支持 NATIVE 格式。
|
> 从 checkpoint 或 savepoint 恢复时,Kafka Source 会优先使用 checkpoint 中保存的 split offset。
diff --git a/docs/zh/connectors/source/ObsFile.md
b/docs/zh/connectors/source/ObsFile.md
index 574501a9c2..349d75af7a 100644
--- a/docs/zh/connectors/source/ObsFile.md
+++ b/docs/zh/connectors/source/ObsFile.md
@@ -75,7 +75,7 @@ import ChangeLog from '../changelog/connector-file-obs.md';
| sheet_name | string | 否 | - |
读取工作簿的工作表,仅在 file_format 为 excel 时使用。
|
| excel_engine | string | 否 | POI | 仅在
`file_format` 为 excel 时使用。支持的引擎包括 `POI` 和 `EasyExcel`。
|
| poi_excel_max_file_size | long | 否 | 52428800 | 仅在
`file_format` 为 excel 且 `excel_engine` 为 POI 时使用。POI 引擎允许读取的最大 Excel 文件大小(默认 50
MB)。
|
-| delimiter | string | 否 | \001 | 字段分隔符
|
+| delimiter/field_delimiter | string | 否 | \001 |
字段分隔符,用于告诉连接器在读取文本文件时如何切分字段。默认 `\001`,与 hive 的默认分隔符相同。**delimiter** 参数将在 2.3.5
版本后废弃,请改用 **field_delimiter**。 |
| row_delimiter | string | 否 | \n | 行分隔符
|
| parse_partition_from_path | boolean | 否 | true |
控制是否从文件路径解析分区键和值 |
| skip_header_row_number | long | 否 | 0 | 跳过前几行,但仅适用于
txt 和 csv。 |
diff --git a/docs/zh/connectors/source/OssFile.md
b/docs/zh/connectors/source/OssFile.md
index d82d7250a3..c99578cf4a 100644
--- a/docs/zh/connectors/source/OssFile.md
+++ b/docs/zh/connectors/source/OssFile.md
@@ -193,7 +193,7 @@ schema {
| read_columns | list | 否 | - |
数据源的读取列列表,用户可以使用它来实现字段投影。支持列投影的文件类型如下所示:`text` `csv` `parquet` `orc` `json`
`excel` `xml`。如果用户想在读取`text` `json` `csv`文件时使用此功能,必须配置"schema"选项。 |
| access_key | string | 否 | - |
|
| access_secret | string | 否 | - |
|
-| delimiter | string | 否 | \001 |
字段分隔符,用于告诉连接器在读取文本文件时如何切分字段。默认`\001`,与hive的默认分隔符相同。
|
+| delimiter/field_delimiter | string | 否 | \001 |
字段分隔符,用于告诉连接器在读取文本文件时如何切分字段。默认`\001`,与hive的默认分隔符相同。**delimiter** 参数将在 2.3.5
版本后废弃,请改用 **field_delimiter**。
|
| row_delimiter | string | 否 | \n |
行分隔符,用于告诉连接器在读取文本文件时如何切分行。默认`\n`。
|
| parse_partition_from_path | boolean | 否 | true |
控制是否从文件路径解析分区键和值。例如,如果您从路径`oss://hadoop-cluster/tmp/seatunnel/parquet/name=tyrantlucifer/age=26`读取文件。文件中的每条记录数据都将添加这两个字段:name="tyrantlucifer",age=16
|
| date_format | string | 否 | yyyy-MM-dd |
日期类型格式,用于告诉连接器如何将字符串转换为日期,支持以下格式:`yyyy-MM-dd` `yyyy.MM.dd`
`yyyy/MM/dd`。默认`yyyy-MM-dd`
|
diff --git a/docs/zh/connectors/source/OssJindoFile.md
b/docs/zh/connectors/source/OssJindoFile.md
index 7f95a361f2..5f0cf631e2 100644
--- a/docs/zh/connectors/source/OssJindoFile.md
+++ b/docs/zh/connectors/source/OssJindoFile.md
@@ -14,7 +14,7 @@ import ChangeLog from
'../changelog/connector-file-oss-jindo.md';
- [x] [批](../../introduction/concepts/connector-v2-features.md)
- [ ] [流](../../introduction/concepts/connector-v2-features.md)
-- [x] [多模态](../../introduction/concepts/connector-v2-features.md#multimodal)
+- [x] [多模态](../../introduction/concepts/connector-v2-features.md#多模态multimodal)
使用二进制文件格式读写任何格式的文件,例如视频、图片等。简而言之,任何文件都可以同步到目标位置。
diff --git a/docs/zh/connectors/source/Persistiq.md
b/docs/zh/connectors/source/Persistiq.md
index 85afbf9178..30d6019d34 100644
--- a/docs/zh/connectors/source/Persistiq.md
+++ b/docs/zh/connectors/source/Persistiq.md
@@ -30,7 +30,7 @@ import ChangeLog from
'../changelog/connector-http-persistiq.md';
| params | Map | 否 | - | HTTP 参数
|
| body | String | 否 | - | HTTP 请求体
|
| json_field | Config | 否 | - | JSON 字段配置
|
-| content_json | String | 否 | - | 内容 JSON 配置
|
+| content_field | String | 否 | - | 内容 JSON 配置
|
| poll_interval_millis | int | 否 | - | 流模式下请求 HTTP API 的间隔(毫秒)
|
| retry | int | 否 | - | 如果 HTTP 请求返回
`IOException` 的最大重试次数 |
| retry_backoff_multiplier_ms | int | 否 | 100 | HTTP 请求失败时的重试退避倍数(毫秒)
|
@@ -84,9 +84,9 @@ HTTP 请求失败时的最大重试退避时间(毫秒)
上游数据的模式字段。更多详情请参考 [Schema 特性](../../introduction/concepts/schema-feature.md)。
-### content_json [String]
+### content_field [String]
-此参数可以获取一些 JSON 数据。
+此参数可以获取一些 JSON 数据。如果只需要 `book` 部分的数据,请配置 `content_field = "$.store.book.*"`。
### json_field [Config]
diff --git a/docs/zh/connectors/source/Typesense.md
b/docs/zh/connectors/source/Typesense.md
index dd412d3a46..1366be9787 100644
--- a/docs/zh/connectors/source/Typesense.md
+++ b/docs/zh/connectors/source/Typesense.md
@@ -48,7 +48,7 @@ Typesense 的访问地址,格式为 `host:port`,例如:`["typesense-01:810
### schema [config]
-typesense
需要读取的列。有关更多信息,请参阅:[guide](../../introduction/concepts/schema-feature.md#how-to-declare-type-supported)。
+typesense
需要读取的列。有关更多信息,请参阅:[guide](../../introduction/concepts/schema-feature.md#如何声明支持的类型)。
### api_key [string]
diff --git a/docs/zh/edge-agent/configuration.md
b/docs/zh/edge-agent/configuration.md
index 0576586402..e79d984e99 100644
--- a/docs/zh/edge-agent/configuration.md
+++ b/docs/zh/edge-agent/configuration.md
@@ -34,7 +34,7 @@ agent.yaml 仅使用顶层配置块:agent、input、queue、retry、output。
| 配置项 | 类型 | 必填 | 默认值 | 说明
|
| -------------------- | ------ | --- | ------------- |
------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|
-| id | string | 否 | 自动生成 | Agent 实例标识(日志、运维)。自动生成见
[身份文件](#身份文件edge-agentid)。
|
+| id | string | 否 | 自动生成 | Agent 实例标识(日志、运维)。自动生成见
[身份文件](#身份文件)。
|
| delivery-guarantee | string | 否 | BEST_EFFORT |
出站投递模式。BEST_EFFORT(别名:best-effort、best_effort):本地 WAL
持久化重试,同一记录可能多次送达,下游需幂等。NON(别名:non、none):无
WAL、无持久化,事件从内存直接发送,失败即丢弃;完全无状态,不依赖本地持久化。详见
[投递模式](./architecture-overview.md#62-投递模式)。 |
| idle-sleep-ms | long | 否 | 200 | 调度循环无进展时的休眠(毫秒)。必须 > 0。
|
| bulk-max-size | integer | 否 | 256 | 内存中待写入 WAL 的事件达到该条数时
flush。 |
@@ -67,7 +67,7 @@ agent.yaml 仅使用顶层配置块:agent、input、queue、retry、output。
| 配置项 | 类型 | 必填 | 默认值 | 说明
|
| ----------------------- | --------- | ----- | -------- |
------------------------------------------------------------- |
-| id | string | 否 | 自动生成 | 采集源标识;WAL/位点
sourceId。自动生成见 [身份文件](#身份文件edge-agentid)。 |
+| id | string | 否 | 自动生成 | 采集源标识;WAL/位点
sourceId。自动生成见 [身份文件](#身份文件)。 |
| type | string | 否 | file | 输入插件类型。当前仅实现 file。
|
| paths | string 列表 | 是 | — | 采集文件 glob(如
/var/log/*.log)。不可为空,不可含空字符串。 |
| encoding | string | 否 | UTF-8 | 文件编码。
|
@@ -164,7 +164,7 @@ endpoint/token 对齐、RAW 与 PACKET 及场景 YAML 见 [输出配置指南](o
| 配置项 | 类型 | 必填 | 默认值 | 说明
|
| ------ | ------ | --- | --------- |
-----------------------------------------------------------------------------------------------------
|
-| id | string | 否 | 自动生成 | 出站逻辑标识(日志、迁移;当前不写入线协议)。自动生成见
[身份文件](#身份文件edge-agentid)。 |
+| id | string | 否 | 自动生成 | 出站逻辑标识(日志、迁移;当前不写入线协议)。自动生成见
[身份文件](#身份文件)。 |
| type | string | 否 | console | transport:EdgeSocket 客户端。console:将 payload 以
EDGE_CONSOLE_OUTPUT 写入 log/edge-agent.log(调试用途)。 |
diff --git a/docs/zh/edge-agent/faq.md b/docs/zh/edge-agent/faq.md
index a5d92751f5..488a3ff54a 100644
--- a/docs/zh/edge-agent/faq.md
+++ b/docs/zh/edge-agent/faq.md
@@ -11,7 +11,7 @@ title: 常见问题
### edge-agent.id 是什么?
-安装根目录下的身份文件(与 edge-agent.pid 同级)。YAML 中省略 agent.id、input.id、output.id 时,Agent
在此读取或写入;YAML 显式 id 优先。详见 [身份文件](configuration.md#身份文件edge-agentid)。
+安装根目录下的身份文件(与 edge-agent.pid 同级)。YAML 中省略 agent.id、input.id、output.id 时,Agent
在此读取或写入;YAML 显式 id 优先。详见 [身份文件](configuration.md#身份文件)。
### agent.id、input.id、output.id 各做什么?
@@ -46,7 +46,7 @@ input.id。位点存在 WAL 数据库的 edge_agent_source_position 表中,按
## WAL 持久化
-默认路径、磁盘文件(data、-wal、-shm)及表内数据说明见[配置说明 — WAL
持久化文件](configuration.md#sqlite-持久化文件)。console 模式仍会使用该库。
+默认路径、磁盘文件(data、-wal、-shm)及表内数据说明见[配置说明 — WAL
持久化文件](configuration.md#wal-持久化文件)。console 模式仍会使用该库。
### WAL 状态 DEAD 是什么意思?
diff --git a/docs/zh/edge-agent/input-configuration.md
b/docs/zh/edge-agent/input-configuration.md
index a9900ebdad..ef8993c485 100644
--- a/docs/zh/edge-agent/input-configuration.md
+++ b/docs/zh/edge-agent/input-configuration.md
@@ -223,7 +223,7 @@ input:
### 8. 固定 input.id
-省略 id 时使用稳定的采集源标识作为 WAL / 位点 sourceId(见
[身份文件](configuration.md#身份文件edge-agentid))。多 Agent 或 WAL 迁移时请保留 edge-agent.id 与
WAL,或显式指定:
+省略 id 时使用稳定的采集源标识作为 WAL / 位点 sourceId(见 [身份文件](configuration.md#身份文件))。多 Agent
或 WAL 迁移时请保留 edge-agent.id 与 WAL,或显式指定:
```yaml
input:
diff --git a/docs/zh/edge-agent/output-configuration.md
b/docs/zh/edge-agent/output-configuration.md
index 861db9d9ea..2de0468bd4 100644
--- a/docs/zh/edge-agent/output-configuration.md
+++ b/docs/zh/edge-agent/output-configuration.md
@@ -175,4 +175,4 @@ output:
## output.id 与迁移
-output.id 用于出站侧日志与迁移标识(不写入线协议)。见
[身份文件](configuration.md#身份文件edge-agentid)。迁移时请一并拷贝 edge-agent.id、WAL 及
config/agent.yaml。
+output.id 用于出站侧日志与迁移标识(不写入线协议)。见 [身份文件](configuration.md#身份文件)。迁移时请一并拷贝
edge-agent.id、WAL 及 config/agent.yaml。
diff --git a/docs/zh/edge-agent/quick-start.md
b/docs/zh/edge-agent/quick-start.md
index 1ebe00ba3c..dbf727535c 100644
--- a/docs/zh/edge-agent/quick-start.md
+++ b/docs/zh/edge-agent/quick-start.md
@@ -10,8 +10,8 @@ title: 快速开始
| 路径 | 需要 Engine | 用途
|
| -------------------------------------------------- | --------- |
----------------------------- |
-| [Console 本地验证](#console-本地验证无需-engine) | 否 | 验证安装、采集、本地 WAL,无需 Engine |
-| [生产模式(接入 Engine)](#生产模式接入-engine) | 是 | Agent 将数据发到 Engine 作业 |
+| [Console 本地验证](#console-本地验证) | 否 | 验证安装、采集、本地 WAL,无需 Engine |
+| [生产模式(接入 Engine)](#生产模式) | 是 | Agent 将数据发到 Engine 作业 |
生产环境加固请参阅 [部署指南](deployment-guide.md) 与 [运维](operations.md)。
@@ -79,7 +79,7 @@ echo '{"event":"world","ts":2}' >>
/tmp/edge-agent-quickstart.log
:::
-仍会生成 edge-agent.id 与 data/ 目录下的 WAL 持久化文件(wal.db、wal.db-wal、wal.db-shm),与是否连接
Engine 无关。说明见[配置说明 — WAL 持久化文件](configuration.md#sqlite-持久化文件)。
+仍会生成 edge-agent.id 与 data/ 目录下的 WAL 持久化文件(wal.db、wal.db-wal、wal.db-shm),与是否连接
Engine 无关。说明见[配置说明 — WAL 持久化文件](configuration.md#wal-持久化文件)。
### 停止
@@ -87,7 +87,7 @@ echo '{"event":"world","ts":2}' >>
/tmp/edge-agent-quickstart.log
sh bin/seatunnel-edge-agent.sh stop
```
-接入 Engine 请参阅 [生产模式](#生产模式接入-engine)。
+接入 Engine 请参阅 [生产模式](#生产模式)。
## 生产模式
diff --git a/docs/zh/engines/zeta/log-analysis-with-ai.md
b/docs/zh/engines/zeta/log-analysis-with-ai.md
index 458cf04e7b..68df847aa8 100644
--- a/docs/zh/engines/zeta/log-analysis-with-ai.md
+++ b/docs/zh/engines/zeta/log-analysis-with-ai.md
@@ -57,7 +57,7 @@ GET http://<master-host>:8080/logs?format=json
GET http://<node-host>:5801/log
```
-第一个接口从所有 Zeta 节点获取匹配的日志,最后一个接口读取单个节点的日志。如果配置了 context path 或动态 HTTP
端口,接口地址也会变化。完整行为请参阅 [RESTful API V2](rest-api-v2.md#get-logs-from-all-nodes)。
+第一个接口从所有 Zeta 节点获取匹配的日志,最后一个接口读取单个节点的日志。如果配置了 context path 或动态 HTTP
端口,接口地址也会变化。完整行为请参阅 [RESTful API V2](rest-api-v2.md#获取所有节点日志内容)。
在 Kubernetes 环境中,需要保留所有相关 Master 和 Worker Pod 的日志。先收集故障时间窗口;如果 Pod
发生过重启,还需要收集上一个容器的日志: