DanielCarter-stack opened a new pull request, #11751:
URL: https://github.com/apache/seatunnel/pull/11751
## Purpose
Improve the connector documentation for GoogleBigtable, GoogleFirestore,
Lance, Typesense and Assert by filling in missing structured metadata,
clarifying the actual behavior of each connector against the current source,
and adding the streaming / multi-table examples that were previously missing.
## Scope of changes
- `docs/en/connectors/source/GoogleBigtable.md` /
`docs/en/connectors/sink/GoogleBigtable.md` (and the matching zh files)
- Add Support These Engines sections and the full Key Features checkbox
list (including column projection / CDC / timer-flush).
- Document that the source is bounded — one split per configured table or
row-key range — and that parallelism does not shard the scan.
- Document the `rowkey_column` STRING vs BYTES requirement and the
lexicographic-only `start_rowkey` / `end_rowkey` behavior.
- Document that every cell mutation is unconditional, so `UPDATE` /
`DELETE` row kinds are not interpreted as CDC operations.
- Add a streaming checkpoint-flush sink example and a bounded streaming
source example that filters by cell versions.
- `docs/en/connectors/sink/GoogleFirestore.md` (and zh)
- Fill out the Key Features checkboxes (stream + timer-flush were missing).
- Clarify that Firestore document IDs are auto-generated, so each row
triggers an `add` call rather than a CDC operation.
- Note the BATCH vs STREAMING behavior (every checkpoint flushes the
in-memory write buffer).
- Add a streaming checkpoint-flush task example.
- `docs/en/connectors/sink/Lance.md` (and zh)
- Fill out the Key Features checkboxes (CDC / batch / stream /
timer-flush).
- Document the current type-mapping caveat: every integer SeaTunnel type
narrows to Arrow `int32`, so `BIGINT` values outside the signed 32-bit range
are truncated; `TIME` uses millisecond precision and `TIMESTAMP` is hardcoded
to Asia/Shanghai.
- Document the `lance.write.enable.stable.row.ids` option (it was listed
in the table but never explained).
- Add APPEND-with-larger-fragment and streaming checkpoint-flush examples.
- `docs/en/connectors/source/Typesense.md` /
`docs/en/connectors/sink/Typesense.md` (and zh)
- Add Support These Engines sections and the full Key Features checkboxes.
- Document that `hosts` is not used for parallel scan sharding on the
source side; the source picks the first reachable node.
- Document that `batch_size` maps to the Typesense `per_page` parameter.
- Document that `UPDATE` / `DELETE` row kinds are upserted by document
`id` (not interpreted as CDC), and add a streaming upsert example paired with
`data_save_mode = DROP_DATA` + a stable `primary_keys`.
- `docs/en/connectors/sink/Assert.md` (and zh)
- Fill out the Key Features checkboxes (CDC / batch / stream /
timer-flush).
- Clarify that Assert is a terminal sink with no external system to write
to, and that `UPDATE` / `DELETE` row kinds are not interpreted as CDC
operations.
- Document the streaming checkpoint-window row-count semantics for
`MIN_ROW` / `MAX_ROW`.
- Add a streaming validation example.
## Notes
- No code or e2e logic was changed; this PR only updates documentation.
- Config names, defaults, and example options match the current connector
implementation.
- I did not bump or edit the supported version metadata for any connector.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]