DanielLeens commented on issue #12127:
URL: https://github.com/apache/seatunnel/issues/12127#issuecomment-5554212730

   Two more sightings today, with server-side detail that narrows the failure 
mode:
   
   - PR #11727, fork run `abdessalems/seatunnel` 33978895542, job 
`all-connectors-it-2 (11)` (job 101340798239): `testFakeSourceToCouchbaseSink` 
failed on the Zeta leg; the Flink 1.18 and 1.20 legs hit the same window but 
recovered through the job's restart.
   - PR #11077, fork run `hesam-oxe/seatunnel` 33971836407, job 
`all-connectors-it-2 (8)` (job 101321961110): failed on the Flink 1.15, Flink 
1.20 and Spark 2.4 legs.
   
   In every case the writer's `waitUntilReady` times out in stage 
`WAIT_FOR_CONFIG` after 30 s, and the KV endpoint keeps failing SASL for the 
whole window:
   
   ```
   [com.couchbase.io][SaslAuthenticationFailedEvent][95ms] Authentication 
Failure - Potential causes: invalid credentials or if LDAP is enabled ensure 
PLAIN SASL mechanism is exclusively used ... 
{"bucket":"test_bucket","remote":"e2e_couchbase:11210","status":"UNKNOWN","type":"KV"}
   [com.couchbase.endpoint][EndpointConnectionFailedEvent][203ms] Connect 
attempt 1 failed because of AuthenticationFailureException ...
   UnambiguousTimeoutException: WaitUntilReady timed out in stage 
WAIT_FOR_CONFIG (spent PT30.001S in that stage) {"bucket":"test_bucket", ... 
"services":{"mgmt":[{"state":"connected", 
"remote":"e2e_couchbase:8091"}],"kv":[{"lastConnectAttemptFailure":"Authentication
 Failure ..."}]}}
   ```
   
   The management port accepts the credentials (`mgmt: connected`), only the KV 
`SELECT_BUCKET` step is refused, and the same credentials work from the test 
JVM (`CouchbaseIT#startUp` already passed `waitUntilReady(2 min)` and ran DDL 
before the job started). The KV service therefore intermittently reports the 
bucket as not selectable for well over 30 s after the IT has verified 
readiness, and the hard-coded `Duration.ofSeconds(30)` in `CouchbaseWriter` 
with no retry turns that into a job failure. The IT cannot pre-empt this from 
the test side: readiness had been confirmed 17-60 s before each failure (the 
Flink 1.18 leg failed 17 s after `Couchbase cluster ready`).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to