This is an automated email from the ASF dual-hosted git repository. github-merge-queue[bot] pushed a commit to branch gh-readonly-queue/dev/pr-12553-0e4cf52896ee3604a8e7899d7736ffd914ee6ce7 in repository https://gitbox.apache.org/repos/asf/seatunnel.git
commit 1ea469f5cc8a2ac5ccce8e7f90b4dc7f3b8bf18a Author: Jast <[email protected]> AuthorDate: Fri Oct 2 07:58:01 2026 +0000 [Docs] Sync outdated connector option tables between English and Chinese docs (#12553) Co-authored-by: zhangshenghang <[email protected]> --- docs/en/connectors/sink/Greenplum.md | 5 ++ docs/en/connectors/source/OssJindoFile.md | 1 + .../connectors/changelog/connector-http-splunk.md | 3 ++ docs/zh/connectors/sink/BosFile.md | 1 + docs/zh/connectors/sink/HdfsFile.md | 7 +++ docs/zh/connectors/sink/ObsFile.md | 1 + docs/zh/connectors/sink/OssFile.md | 1 + docs/zh/connectors/sink/S3File.md | 6 +++ docs/zh/connectors/sink/SftpFile.md | 1 + docs/zh/connectors/sink/StarRocks.md | 1 + docs/zh/connectors/source/BosFile.md | 1 + docs/zh/connectors/source/PostgreSQL.md | 3 +- docs/zh/connectors/source/Splunk.md | 54 ++++++++++++++++++++++ 13 files changed, 84 insertions(+), 1 deletion(-) diff --git a/docs/en/connectors/sink/Greenplum.md b/docs/en/connectors/sink/Greenplum.md index 72df9e6ad3..088086c589 100644 --- a/docs/en/connectors/sink/Greenplum.md +++ b/docs/en/connectors/sink/Greenplum.md @@ -62,6 +62,11 @@ Only Greenplum-specific commonly used options are listed here. Other JDBC sink o | generate_sink_sql | Boolean | No | false | Generate the insert SQL automatically from `database` and `table`. When `true`, the column order must match the upstream schema. | | database | String | No | - | Database name used when `generate_sink_sql = true`. | | table | String | No | - | Target table name used when `generate_sink_sql = true`. | +| primary_keys | Array | No | - | Primary key fields used for upsert semantics when the sink SQL is generated automatically. | +| connection_check_timeout_sec | Int | No | 30 | The time (seconds) to wait for database operation used to validate the connection. | +| max_commit_attempts | Int | No | 3 | Retry times when a transaction commit fails. | +| transaction_timeout_sec | Int | No | -1 | Transaction timeout in seconds. `-1` means unlimited. | +| enable_upsert | Boolean | No | true | Enable upsert write by primary keys. | | common-options | | No | - | Sink plugin common parameters, please refer to [Sink Common Options](../common-options/sink-common-options.md) for details. | :::tip diff --git a/docs/en/connectors/source/OssJindoFile.md b/docs/en/connectors/source/OssJindoFile.md index 7ad4d49fd1..202f85d598 100644 --- a/docs/en/connectors/source/OssJindoFile.md +++ b/docs/en/connectors/source/OssJindoFile.md @@ -82,6 +82,7 @@ It only supports hadoop version **2.9.X+**. | xml_use_attr_format | boolean | no | - | | csv_use_header_line | boolean | no | false | | file_filter_pattern | string | no | | +| filename_extension | string | no | - | | compress_codec | string | no | none | | archive_compress_codec | string | no | none | | encoding | string | no | UTF-8 | diff --git a/docs/zh/connectors/changelog/connector-http-splunk.md b/docs/zh/connectors/changelog/connector-http-splunk.md new file mode 100644 index 0000000000..c2e6a88eea --- /dev/null +++ b/docs/zh/connectors/changelog/connector-http-splunk.md @@ -0,0 +1,3 @@ +| Change | Commit | Version | +| --- | --- | --- | +| [Feature][connector-http-splunk] Add Splunk source connector | https://github.com/apache/seatunnel | dev | \ No newline at end of file diff --git a/docs/zh/connectors/sink/BosFile.md b/docs/zh/connectors/sink/BosFile.md index a47f2eed55..1540a553cd 100644 --- a/docs/zh/connectors/sink/BosFile.md +++ b/docs/zh/connectors/sink/BosFile.md @@ -84,6 +84,7 @@ import ChangeLog from '../changelog/connector-file-bos.md'; | parquet_avro_write_timestamp_as_int96 | boolean | 否 | false | parquet 格式使用 | | parquet_avro_write_fixed_as_int96 | array | 否 | - | parquet 格式使用 | | encoding | string | 否 | "UTF-8" | json/text/csv/xml 格式使用 | +| common-options | object | 否 | - | Sink 插件通用参数,请参考 [Sink Common Options](../common-options/sink-common-options.md) | ## 示例 diff --git a/docs/zh/connectors/sink/HdfsFile.md b/docs/zh/connectors/sink/HdfsFile.md index d60909c12b..bb143afcb5 100644 --- a/docs/zh/connectors/sink/HdfsFile.md +++ b/docs/zh/connectors/sink/HdfsFile.md @@ -75,8 +75,15 @@ import ChangeLog from '../changelog/connector-file-hadoop.md'; | kerberos_keytab_path | string | 否 | - | kerberos 的 keytab 路径 | | common-options | object | 否 | - | 接收器插件通用参数,请参阅 [接收器通用选项](../common-options/sink-common-options.md) 了解详情 | | csv_string_quote_mode | enum | 否 | MINIMAL | 仅在文件格式为 CSV 时使用。 | +| xml_root_tag | string | 否 | RECORDS | 仅在 file_format 为 xml 时使用,指定 XML 文件中根元素的标签名称。 | +| xml_row_tag | string | 否 | RECORD | 仅在 file_format 为 xml 时使用,指定 XML 文件中数据行的标签名称。 | +| xml_use_attr_format | boolean | 否 | - | 仅在 file_format 为 xml 时使用,指定是否使用标签属性格式处理数据。 | | enable_header_write | boolean | 否 | false | 仅在 file_format_type 为 text,csv 时使用。<br/> false:不写入表头,true:写入表头。 | +| parquet_avro_write_timestamp_as_int96 | boolean | 否 | false | 仅在 file_format 为 parquet 时使用。 | +| parquet_avro_write_fixed_as_int96 | array | 否 | - | 仅在 file_format 为 parquet 时使用。 | +| encoding | string | 否 | "UTF-8" | 仅在 file_format_type 为 json,text,csv,xml 时使用。 | | max_rows_in_memory | int | 否 | - | 仅当 file_format 为 excel 时使用。当文件格式为 Excel 时,可以缓存在内存中的最大数据项数。 | +| sheet_max_rows | int | 否 | 1048576 | 仅在 file_format 为 excel 时使用,每个工作表允许写入的最大行数。 | | sheet_name | string | 否 | Sheet0 | 仅当 file_format 为 excel 时使用。将工作簿的表写入指定的表名 | | remote_user | string | 否 | - | Hdfs的远端用户名。 | | schema_evolution_enabled | boolean | 否 | false | 开启 Schema 演变支持,适用于 CDC 管道。为 true 时,来自上游的 ADD/DROP/RENAME/MODIFY 列事件无需重启作业即可应用到 Sink。不支持 binary 格式。 | diff --git a/docs/zh/connectors/sink/ObsFile.md b/docs/zh/connectors/sink/ObsFile.md index 484f579c45..4d40ff93c6 100644 --- a/docs/zh/connectors/sink/ObsFile.md +++ b/docs/zh/connectors/sink/ObsFile.md @@ -88,6 +88,7 @@ import ChangeLog from '../changelog/connector-file-obs.md'; | compress_codec | string | 否 | none | 文件的压缩编解码器。Excel 格式不支持任何压缩格式。[提示](#compress_codec) | | common-options | object | 否 | - | Sink 插件通用参数,请参考 [Sink Common Options](../common-options/sink-common-options.md)。[提示](#common_options) | | max_rows_in_memory | int | 否 | - | 当文件格式为Excel时,内存中可以缓存的最大数据项数。仅在file_format_type为excel时使用。 | +| sheet_max_rows | int | 否 | 1048576 | 仅在 file_format 为 excel 时使用,每个工作表允许写入的最大行数。 | | sheet_name | string | 否 | Sheet0 | 标签页。仅在file_format_type为excel时使用。 | | merge_update_event | boolean | 否 | false | 仅当file_format_type为canal_json、debezium_json、maxwell_json 时使用。设置为 `true` 时,会将 `UPDATE_AFTER` 与 `UPDATE_BEFORE` 合并为 `UPDATE` 事件数据。 | | schema_evolution_enabled | boolean | 否 | false | 开启 Schema 演变支持,适用于 CDC 管道。为 true 时,来自上游的 ADD/DROP/RENAME/MODIFY 列事件无需重启作业即可应用到 Sink。不支持 binary 格式。 | diff --git a/docs/zh/connectors/sink/OssFile.md b/docs/zh/connectors/sink/OssFile.md index 1f1928c5db..f43caf99b8 100644 --- a/docs/zh/connectors/sink/OssFile.md +++ b/docs/zh/connectors/sink/OssFile.md @@ -109,6 +109,7 @@ import ChangeLog from '../changelog/connector-file-oss.md'; | file_name_expression | string | 否 | "${transactionId}" | 仅在custom_filename为true时使用 | | filename_time_format | string | 否 | "yyyy.MM.dd" | 仅在custom_filename为true时使用 | | file_format_type | string | 否 | "csv" | 文件格式类型,支持:`text`、`csv`、`parquet`、`orc`、`json`、`excel`、`xml`、`binary`、`canal_json`、`debezium_json`、`maxwell_json`。 | +| filename_extension | string | 否 | - | 使用自定义的文件扩展名覆盖默认的文件扩展名。例如:`.xml`、`.json`、`dat`、`.customtype`。 | | field_delimiter | string | 否 | '\001' for text and ',' for csv | 仅当file_format_type为文本时使用 | | row_delimiter | string | 否 | "\n" | 仅当file_format_type为 `text`、`csv`、`json` 时使用 | | have_partition | boolean | 否 | false | 是否需要处理分区。 | diff --git a/docs/zh/connectors/sink/S3File.md b/docs/zh/connectors/sink/S3File.md index f334845c98..d1990c51bc 100644 --- a/docs/zh/connectors/sink/S3File.md +++ b/docs/zh/connectors/sink/S3File.md @@ -116,6 +116,7 @@ import ChangeLog from '../changelog/connector-file-s3.md'; | file_name_expression | string | 否 | "${transactionId}" | 仅当 custom_filename 为 true 时使用 | | filename_time_format | string | 否 | "yyyy.MM.dd" | 仅当 custom_filename 为 true 时使用 | | file_format_type | string | 否 | "csv" | | +| filename_extension | string | 否 | - | 使用自定义的文件扩展名覆盖默认的文件扩展名。例如:`.xml`、`.json`、`dat`、`.customtype` | | field_delimiter | string | 否 | '\001' for text and ',' for csv | 仅当 file_format 为 text 时使用 | | row_delimiter | string | 否 | "\n" | 仅当 file_format 为 `text`、`csv`、`json` 时使用 | | have_partition | boolean | 否 | false | 是否需要处理分区。 | @@ -128,6 +129,7 @@ import ChangeLog from '../changelog/connector-file-s3.md'; | compress_codec | string | 否 | none | | | common-options | object | 否 | - | | | max_rows_in_memory | int | 否 | - | 仅当 file_format 为 excel 时使用 | +| sheet_max_rows | int | 否 | 1048576 | 仅当 file_format 为 excel 时使用,每个工作表允许写入的最大行数。 | | sheet_name | string | 否 | Sheet0 | 仅当 file_format 为 excel 时使用 | | csv_string_quote_mode | enum | 否 | MINIMAL | 仅当 file_format 为 csv 时使用 | | xml_root_tag | string | 否 | RECORDS | 仅当 file_format 为 xml 时使用,指定 XML 文件中根元素的标签名称。 | @@ -267,6 +269,10 @@ Sink 插件通用参数,请参考 [Sink 通用选项](../common-options/sink-c 当文件格式为 Excel 时,内存中可以缓存的最大数据项数。 +### sheet_max_rows [int] + +当文件格式为 Excel 时,每个工作表允许写入的最大行数。 + ### sheet_name [string] 写入工作表的名称 diff --git a/docs/zh/connectors/sink/SftpFile.md b/docs/zh/connectors/sink/SftpFile.md index b947cd0e20..7eaa932ef0 100644 --- a/docs/zh/connectors/sink/SftpFile.md +++ b/docs/zh/connectors/sink/SftpFile.md @@ -57,6 +57,7 @@ import ChangeLog from '../changelog/connector-file-sftp.md'; | file_name_expression | string | 否 | "${transactionId}" | 仅在custom_filename为true时使用 | | filename_time_format | string | 否 | "yyyy.MM.dd" | 仅在custom_filename为true时使用 | | file_format_type | string | 否 | "csv" | | +| filename_extension | string | 否 | - | 使用自定义的文件扩展名覆盖默认的文件扩展名。例如:`.xml`、`.json`、`dat`、`.customtype` | | field_delimiter | string | 否 | '\001' for text and ',' for csv | 仅当file_format_type为text时使用 | | row_delimiter | string | 否 | "\n" | 仅当file_format_type为 `text`、`csv`、`json` 时使用 | | have_partition | boolean | 否 | false | 是否需要处理分区。 | diff --git a/docs/zh/connectors/sink/StarRocks.md b/docs/zh/connectors/sink/StarRocks.md index fb9a99efd4..1cf935b7d3 100644 --- a/docs/zh/connectors/sink/StarRocks.md +++ b/docs/zh/connectors/sink/StarRocks.md @@ -54,6 +54,7 @@ StarRocks数据接收器内部实现采用了缓存,通过stream load将数据 | http_socket_timeout_ms | int | 否 | 180000 | HTTP socket 超时时间,默认为 3 分钟 | | schema_save_mode | Enum | 否 | CREATE_SCHEMA_WHEN_NOT_EXIST | 同步任务启动前,针对目标端已存在的表结构选择不同处理方式 | | data_save_mode | Enum | 否 | APPEND_DATA | 同步任务启动前,针对目标端已存在的数据选择不同处理方式 | +| table_options | Map | 否 | - | SaveMode 自动建表时合并进 CREATE TABLE PROPERTIES 的 Sink 专属表属性,详见表下方说明 | | custom_sql | String | 否 | - | 当 `data_save_mode` 设置为 `CUSTOM_PROCESSING` 时必须配置。该 SQL 会在同步任务启动前执行 | ### save_mode_create_template diff --git a/docs/zh/connectors/source/BosFile.md b/docs/zh/connectors/source/BosFile.md index de61486751..f6d76d290d 100644 --- a/docs/zh/connectors/source/BosFile.md +++ b/docs/zh/connectors/source/BosFile.md @@ -90,6 +90,7 @@ import ChangeLog from '../changelog/connector-file-bos.md'; | escape_char | string | 否 | - | | recursive_file_scan | boolean | 否 | true | | sort_files_by_modification_time | boolean | 否 | false | +| common-options | | 否 | - | ## 示例 diff --git a/docs/zh/connectors/source/PostgreSQL.md b/docs/zh/connectors/source/PostgreSQL.md index 81c4eecdda..5d8bea7a63 100644 --- a/docs/zh/connectors/source/PostgreSQL.md +++ b/docs/zh/connectors/source/PostgreSQL.md @@ -103,7 +103,8 @@ import ChangeLog from '../changelog/connector-jdbc.md'; | split.even-distribution.factor.upper-bound | Double | 否 | 100 | 块键分布因子的上限。此因子用于确定表数据是否均匀分布。<br/> 如果计算出的分布因子小于或等于此上限(即 (MAX(id) - MIN(id) + 1) / 行数),则表块将优化为均匀分布。否则,如果分布因子较大,则将视为不均匀分布,当估计的分片数超过 `sample-sharding.threshold` 指定的值时,将使用基于采样的分片策略。默认值为 100.0。 | | split.sample-sharding.threshold | Int | 否 | 1000 | 此配置指定触发样本分片策略的估计分片数阈值。<br/> 当分布因子超出 `chunk-key.even-distribution.factor.upper-bound` 和 `chunk-key.even-distribution.factor.lower-bound` 指定的范围时,且估计的分片数(计算为近似行数 / 块大小)超过此阈值,将使用样本分片策略。这可以帮助更高效地处理大数据集。默认值为 1000 个分片。 | | split.inverse-sampling.rate | Int | 否 | 1000 | 在样本分片策略中使用的采样率的逆数。例如,如果此值设置为 1000,表示在采样过程中应用 1/1000 的采样率。此选项提供了控制采样粒度的灵活性,从而影响最终的分片数量。在处理非常大的数据集时,较低的采样率尤其有用。默认值为 1000。 | -| +| common-options | | 否 | - | Source 插件通用参数,请参阅 [Source Common Options](../common-options/source-common-options.md) 了解详情 [...] + ## 并行读取器 JDBC 源连接器支持从表中并行读取数据。SeaTunnel 将使用某些规则来拆分表中的数据,这些数据将交给读取器进行读取。读取器的数量由 `parallelism` 选项确定。 diff --git a/docs/zh/connectors/source/Splunk.md b/docs/zh/connectors/source/Splunk.md new file mode 100644 index 0000000000..758fcdb947 --- /dev/null +++ b/docs/zh/connectors/source/Splunk.md @@ -0,0 +1,54 @@ +import ChangeLog from '../changelog/connector-http-splunk.md'; + +# Http-Splunk Source Connector + +## 描述 + +`Http-Splunk` 连接器允许通过 HTTP Source 模式批量读取 Splunk REST API 端点的数据。它使用基于 Token 的认证,并支持通过 POST 请求访问同步搜索导出端点。 + +## 主要特性 + +* 基于 HTTP 从 Splunk REST API(`/services/search/v2/jobs/export`)摄取数据 +* 通过 `Authorization` 请求头进行基于 Token 的认证 +* 使用表单编码(form-urlencoded)参数映射搜索查询与输出格式 +* 内存保护机制(`max_response_size_bytes`),在大数据量导出时保护 Worker 堆内存 + +## 选项 + +| 名称 | 类型 | 必填 | 默认值 | 描述 | +| --- | --- | --- | --- | --- | +| url | String | 是 | - | Splunk REST API 搜索导出 URL(`/services/search/v2/jobs/export`) | +| api_key | String | 是 | - | Splunk 认证 Token(例如 `Splunk <token>`) | +| method | String | 否 | `POST` | HTTP 请求方法 | +| keep_params_as_form | Boolean | 否 | `true` | 以表单编码(form urlencoded)形式传递参数 | +| max_response_size_bytes | Long | 否 | `20971520` | 允许的最大 HTTP 响应大小(字节),超过后快速失败,以避免大数据量导出时内存无限增长。 | +| params | Map | 是 | - | 请求参数,包括 `search` 和 `output_mode` | + +## 示例配置 + +```hocon +source { + Splunk { + url = "https://your-splunk-instance:8089/services/search/v2/jobs/export" + api_key = "Splunk your_splunk_auth_token" + method = "POST" + keep_params_as_form = true + max_response_size_bytes = 20971520 + params { + search = "search index=_internal | head 10" + output_mode = "json" + } + plugin_output = "splunk_data" + } +} +``` + +## 大数据量导出与内存注意事项 + +Splunk 的导出端点可能返回任意大的结果集。由于该连接器会将完整响应缓冲在内存中,并通过中间字符串表示进行解析,导出时的峰值内存占用约为原始响应大小的数倍(约 2x–3x)。 + +* **默认限制:** `max_response_size_bytes` 选项默认为 **20MB**(`20971520` 字节)。 +* **Worker 堆内存规划:** 请确保 Worker 容器/JVM 堆内存相对于该限制进行了合理的规划。如果您的搜索返回大量数据,请使用 Splunk 的 `earliest`/`latest` 参数缩小搜索时间窗口,或使用更小的 `head` 限制。 +* **调优:** 只有在确认集群 Worker 有足够的堆内存余量可以安全承受更大导出的内存放大倍数后,才调整 `max_response_size_bytes`。 + +<ChangeLog />
