RaphaelAdewale opened a new pull request, #11838:
URL: https://github.com/apache/seatunnel/pull/11838

   ### Purpose of this pull request
   
   Implements the **Splunk Sink** connector, the first half of #10753. The 
Source
   will follow in a separate PR.
   
   The sink writes SeaTunnel rows to a Splunk index through the HTTP Event
   Collector. Each row becomes one HEC event envelope — the row under `event`,
   with `index` / `source` / `sourcetype` / `host` / `time` populated from sink
   options — and envelopes are POSTed to `/services/collector/event` in batches.
   
   Part of #10753
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes — a new `Splunk` sink plugin. No existing option, default, or API is
   changed. Docs added at `docs/en/connectors/sink/Splunk.md` and the zh
   equivalent.
   
   ### Design notes
   
   - **Not built on `connector-http-base`.** Its `HttpClientProvider` always
     builds `HttpClients.createDefault()` with no TLS hook, and Splunk HEC
     commonly sits behind the self-signed certificate generated at install time.
     `SplunkHecClient` adds `tls_verify_certificate` / `tls_verify_hostname`
     (named to match the Elasticsearch sink).
   - **Retry classification.** Only transport errors, HTTP 429 and 5xx are
     retried. A bad token, a forbidden index, or a malformed payload fails the
     task immediately instead of burning retries on an error that cannot resolve
     itself. The buffer is cleared only after the collector accepts the batch, 
so
     a failed attempt never silently drops events.
   - **No connector-level `flush_interval`.** Periodic flushing uses the
     engine-level `sink.flush.interval` via `context.registerFlushAction`,
     consistent with the direction taken in connector-prometheus and the
     Elasticsearch/Hudi sinks. Zeta only; Spark and Flink flush on
     `max_batch_size`, checkpoint and close.
   - **`time_field` accepts** `TIMESTAMP` (read as UTC), `TIMESTAMP_TZ`, and
     `BIGINT` (epoch millis). Any other type fails at startup with a message
     naming the field and its type. The value is sent as epoch seconds with
     millisecond precision, written in plain notation — Splunk rejects the
     scientific notation Jackson produces for doubles at epoch magnitude.
   - **Delivery is at-least-once** and documented as such; HEC offers no
     server-side dedup for this endpoint.
   
   ### How was this patch tested?
   
   - 34 unit tests: config validation and endpoint normalisation, HEC envelope
     serialization (metadata mapping, timestamp conversion, plain-notation wire
     format), batching and flush thresholds, retry classification, and the
     factory OptionRule.
   - E2E `SplunkIT`: `splunk/splunk:9.3.2` container, FakeSource → Splunk sink,
     asserting the events land in the target index with the configured metadata.
   - Additionally verified by hand against a real Splunk 9.3.2 instance using 
the
     built distribution: a Zeta job wrote 5 rows, and a Splunk search confirmed
     the event bodies, `host` from `host_field`, static `source`/`sourcetype`,
     and `_time` from `time_field`. An invalid-token run confirmed the fail-fast
     path.
   
   Note: the local Docker version (29.3.0) is incompatible with this repo's
   pinned Testcontainers client, so no E2E in the repo can run locally — the
   pre-existing `HttpIT` fails identically. `SplunkIT` therefore runs for the
   first time in CI; the manual round-trip above covers the same ground.
   
   ### Check list
   
   * [x] Code changed are covered with tests
   * [x] If any new Jar binary package adding in your PR, please add License 
Notice according [New License 
Guide](https://github.com/apache/seatunnel/blob/dev/docs/en/contribution/new-license.md)
 — no new third-party dependency; `httpclient`/`httpcore` are already in 
`known-dependencies.txt`
   * [x] If necessary, please update the documentation to describe the new 
feature
   * [x] Update the `plugin-mapping.properties` and add new connector 
information in it
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to