davidzollo commented on PR #10447:
URL: https://github.com/apache/seatunnel/pull/10447#issuecomment-5526745762

   Pushed a fix for the `Build` CI failure on the previous head (`5ce9695c5`): 
`fbeb29253`.
   
   **Root cause**: `updated-modules-integration-test-part-1` (JDK 8 and JDK 11, 
all engines — Zeta and Flink 1.13/1.15/1.18/1.20) was failing deterministically 
with `ASSERT-01: row num :2 fail rule: MIN_ROW=4.0` in 
`HubSpotIT#testHubspotSourcePagination`. I downloaded and read the raw job 
logs: MockServer *did* correctly serve both pages (I can see both HTTP requests 
land — the first with no `after` param, the second with `after=token-123` — and 
MockServer returns the right body for each, including page 2's Charlie/David). 
So cursor pagination itself (the thing this PR/E2E test was actually built to 
verify) is working correctly. The bug is one level deeper, in the shared 
`connector-http-base` reader that both requests flow through.
   
   `HttpSourceReader` reuses one `jsonConfiguration` (with 
`Option.ALWAYS_RETURN_LIST`) for `content_field` extraction, `json_field` 
extraction, and cursor extraction. For a *definite* JsonPath (no wildcards) — 
which `content_field` always is — the Jayway JsonPath library wraps the 
already-resolved value in a synthetic single-element list before returning it. 
So `content_field = "$.results"` against a response whose `results` is itself a 
2-element array comes back double-nested as `[[obj1, obj2]]` instead of `[obj1, 
obj2]`. `JsonDeserializationSchema#collect` only unwraps one array level, so 
the *entire inner array* collapsed into a single malformed row instead of 
splitting into 2 — silently turning 4 total rows (2 pages × 2) into 2.
   
   `content_field` isn't set by any other connector or test in the repo 
(verified with a full-repo grep), so this shared-code bug was simply never 
exercised before this PR's E2E test.
   
   **Fix** (in `HttpSourceReader.getPartOfJson()`): unwrap the synthetic 
definite-path wrapper before handing content to the deserializer, mirroring the 
JsonPath library's own contract exactly — only *definite* paths get unwrapped; 
an indefinite/wildcard `content_field` (which never gets the synthetic wrap) is 
left untouched. The shared `jsonConfiguration` field itself isn't touched, so 
`json_field` and cursor extraction behavior for every other HTTP-based 
connector is unaffected.
   
   Also added 
`HttpSourceReaderInternalPollNextTest#testDefiniteArrayContentFieldProducesOneRowPerElement`,
 which drives the real `HttpSourceReader → DeserializationCollector → 
JsonDeserializationSchema` pipeline with a 2-element `content_field` array and 
asserts both elements come back as separate rows — this is the regression guard 
for the exact bug the HubSpot E2E job hit.
   
   Verified locally with `./mvnw spotless:apply -pl 
seatunnel-connectors-v2/connector-http/connector-http-base -am -nsu` (no 
formatting changes needed). Compilation, the new unit test, and the HubSpot E2E 
job itself will be validated by CI on this head.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to