zhangshenghang opened a new pull request, #12499:
URL: https://github.com/apache/seatunnel/pull/12499

   ## Purpose of the PR
   
   This PR fixes a batch of real documentation issues in the file-family, 
messaging and database connector docs, all verified against the actual 
connector source code on `dev`. It covers 50 documentation files (EN and ZH 
together).
   
   ## What was found and changed
   
   All of the following were verified against the corresponding `*Options.java` 
/ factory `OptionRule` before editing:
   
   ### Incorrect values (doc vs code)
   
   - `sheet_name` sink default was documented as `Sheet${Random number}`, but 
the code always names the first sheet `Sheet0` (`ExcelGenerator.java`). Fixed 
in 10 EN + 10 ZH sink docs (HdfsFile, LocalFile, S3File, CosFile, ObsFile, 
OssJindoFile, OssFile, SftpFile, FtpFile, SmbFile).
   - `tmp_path` was marked **required** in HdfsFile / FtpFile / SftpFile sink 
docs, but the code declares it optional with default `/tmp/seatunnel` 
(`FileBaseSinkOptions`, factory `.optional(...)`). Fixed in EN + ZH.
   - Sink `field_delimiter` default cell said only `'\001'`, but for `csv` the 
effective default is `,` (`BaseFileSinkConfig`). Updated the default cell in 9 
EN + 9 ZH sink docs.
   - InfluxDB sink `multi_table_sink_replica` default shown as `-`, actual 
default is `1` (`SinkConnectorCommonOptions`). Fixed EN + ZH.
   - Socket sink docs claimed `max_retries = -1` enables infinite retry, but 
the option rule validates `>= 0` (`SocketSinkFactory` uses 
`Conditions.greaterOrEqual`), so `-1` can never be used. Reworded table row and 
FAQ in EN + ZH.
   - Milvus source ZH doc described `rate_limit` as a client-side read-rate 
hint; the code mutates the collection-wide Milvus property 
`collection.queryRate.max.qps`, which affects all clients of the collection. ZH 
now matches the (already correct) EN description and notes.
   - MongoDB source ZH tip claimed `match.query` and legacy `matchQuery` 
"cannot be set together"; the code treats `matchQuery` as a fallback key and 
`match.query` wins. ZH now matches EN.
   
   ### Malformed tables / duplicated content
   
   - HdfsFile sink (EN+ZH) had a duplicated short `compress_codec` row; removed.
   - S3File source EN had a duplicated `csv_use_header_line` row; removed.
   - HdfsFile source (EN+ZH) contained the whole "Incremental Sync" + 
"Continuous Discovery" example sections duplicated verbatim (78 lines each); 
removed the second copies.
   - Fixed pre-existing malformed option-table rows that rendered with a 
dropped or missing description cell: zh CosFile, zh FtpFile, zh LocalFile 
(extra cell in 4-column tables), zh S3File + zh SftpFile `row_delimiter` 
(trailing empty cell), SftpFile EN+ZH `archive_compress_codec`/`encoding` and 
HdfsFile EN+ZH `archive_compress_codec` (empty description cells — descriptions 
added).
   - FtpFile source EN `csv_use_header_line` default `-` corrected to `false` 
(ZH already said `false`).
   
   ### Missing option rows (options exist in code / shared read path but absent 
from docs)
   
   - Added `read_partitions` (defined in `FileBaseOptions`, honored by the 
shared `AbstractReadStrategy`) to all 12 file source docs, EN + ZH, matching 
each table's column layout.
   - ObsFile source (EN+ZH) was missing 12 shared options that the read path 
honors (`xml_row_tag`, `xml_use_attr_format`, `csv_use_header_line`, 
`compress_codec`, `archive_compress_codec`, `encoding`, `null_format`, 
`binary_chunk_size`, `binary_complete_file_mode`, `file_filter_pattern`, 
`enable_file_split`, `file_split_size`); rows added from sibling docs. Note: 
the continuous-discovery/sync option family was intentionally *not* added 
because `ObsFileSource` extends `BaseFileSource`, unlike `OssFileSource` which 
extends `BaseMultipleTableFileSource`.
   - S3File source (EN+ZH) was missing `file_filter_modified_start` / 
`file_filter_modified_end`; rows added.
   - SmbFile source (EN+ZH): key-features list contradicted the 
`file_format_type` row and examples (parquet/orc claimed by the table but 
missing from the feature checklist). Since `FileFormat.PARQUET/ORC` map to the 
filesystem-agnostic `ParquetReadStrategy`/`OrcReadStrategy`, parquet and orc 
were added to the checklist.
   - SmbFile sink (EN+ZH) was missing `multi_table_sink_replica` although 
`SmbFileSinkFactory` declares it; row added.
   
   ## Language check
   
   Yes — every issue was fixed in **both** `docs/en` and `docs/zh` where both 
versions exist.
   
   ## Duplicate PR check
   
   Before opening this PR I checked all PRs opened in the last 7 days (both 
open and merged): #12487, #12476, #12459, #12423, #12447, #12409 and collected 
their full changed-file lists. None of the files or issues touched by this PR 
overlap with them: those PRs covered transforms, Zeta engine docs, 
Doris/SelectDB and other connector docs (Jdbc, Kafka source, MySQL CDC, Paimon, 
etc.), while this PR targets the file-family connectors 
(Hdfs/Local/S3/OSS/OSS-Jindo/COS/OBS/BOS/FTP/SFTP/SMB/GCS), ObsFile/SmbFile 
gaps, plus small InfluxDB/Socket/Milvus/MongoDB fixes that no recent PR 
includes.
   
   ## Verification
   
   - Every changed default/required flag was read back from the corresponding 
Java options/factory before editing (paths cited above).
   - A structural check was run over all 50 modified files: every markdown 
option table now has consistent cell counts per row (0 malformed rows; the 
script also caught and fixed pre-existing malformed rows).
   - All relative `.md` links in the 50 modified files were checked to resolve 
(0 broken).
   - A full Docusaurus build was not run because the website toolchain lives in 
the separate `apache/seatunnel-website` repository; the checks above cover the 
same failure modes (broken links, malformed tables) at lower cost. No Java code 
was changed, so `spotless`/build verification is not applicable.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to