FrankChen021 commented on issue #19036:
URL: https://github.com/apache/druid/issues/19036#issuecomment-5750076421

   I investigated this against the current code. Disclosure: I’m Codex, an AI 
coding agent from OpenAI, posting this analysis on behalf of the repository 
user.
   
   This looks like a Web Console sampling regression rather than a 
`quantilesDoublesSketch` ingestion problem, and it applies to complex metric 
types generally (consistent with the reported `doubleLast` case).
   
   The classic loader's Connect step sends a sampler schema with:
   
   ```json
   "dimensionsSpec": {
     "useSchemaDiscovery": true
   }
   ```
   
   For a `druid` input source, `DruidSegmentReader` returns complex metric 
values as their native Java objects. Since the Connect request has no 
`metricsSpec` or dimension exclusions yet, schema discovery treats every 
non-time field as a dimension. The sampler then adds the row to its temporary 
incremental index, where `AutoTypeColumnIndexer` receives the sketch as 
expression type `COMPLEX` and throws `Unhandled type: COMPLEX`.
   
   This also explains why manually submitting a spec with the correct 
aggregator works: the complex field is then handled as a metric instead of 
being passed to the auto-type dimension indexer.
   
   The likely regression is #17160, which changed the Connect sampler's 
`dimensionsSpec` from `{}` to `{ useSchemaDiscovery: true }` for fixed-format 
sources to support Delta, but applied the change to Druid input sources as 
well. The problematic behavior is still present on current `master`.
   
   A narrow fix in `sampleForConnect` would be to retain type-aware discovery 
for other sources but restore the old dimensions spec for reingestion:
   
   ```ts
   dimensionsSpec: reingestMode
     ? {}
     : {
         useSchemaDiscovery: true,
         dimensions: addFileUri ? ['__file_uri'] : undefined,
       },
   ```
   
   This allows the initial Druid sample to complete. The existing `scan` and 
`segmentMetadata` queries can then populate the actual dimensions and 
`metricsSpec`, as they already do after the sample succeeds. I would avoid 
changing `AutoTypeColumnIndexer` to accept arbitrary complex objects because 
complex values are not valid auto-typed/nested dimension values and that would 
have much broader semantics.
   
   A regression test should cover a Druid Connect sample containing a complex 
metric, plus ensure Delta and other sources retain `useSchemaDiscovery: true`. 
The temporary-segment cleanup warning appears secondary to the sampler abort, 
not the primary problem.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to