[ 
https://issues.apache.org/jira/browse/CAMEL-25027?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jiri Ondrusek updated CAMEL-25027:
----------------------------------
    Summary: camel-langchain4j-ingest - Add media modality to embed images, 
audio, video and PDF whole     (was: camel-langchain4j-ingest - Add audio 
modality to embed audio documents whole)

> camel-langchain4j-ingest - Add media modality to embed images, audio, video 
> and PDF whole   
> --------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25027
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25027
>             Project: Camel
>          Issue Type: New Feature
>          Components: camel-langchain4j
>    Affects Versions: 4.23.0
>            Reporter: Jiri Ondrusek
>            Assignee: Jiri Ondrusek
>            Priority: Major
>
> The langchain4j-ingest producer only handles text: the body is read as a 
> String, split into segments and embedded with embedAll. Audio files cannot be 
> ingested as audio
>   embeddings, the way a Wav2Vec2 or CLAP model would embed them for acoustic 
> similarity search. Without a parser they are decoded as garbage text and 
> embedded anyway.
>   LangChain4j 1.19 has the API for this: EmbeddingRequest accepts 
> AudioContent and a model declares AUDIO in supportedContentTypes().
>  
>   Proposal: a modality endpoint option, text (default, unchanged behaviour) 
> or audio.
>   With modality=audio:
>   - the body is read as bytes and embedded whole, as one vector, through 
> embed(EmbeddingRequest) with an AudioContent
>   - the vector is stored with a placeholder TextSegment whose text is the 
> document id, carrying the same camel_ingest_pipeline and 
> camel_ingest_document_id metadata as text
>     segments
>   - the endpoint fails to start when the model does not declare AUDIO
>   - the MIME type comes from a new contentType option or, when unset, from 
> the document id's file extension (wav, mp3, flac, ogg, m4a, aac); an unknown 
> type fails the
>     exchange
>   - maxDocumentSize and minDocumentSize count bytes; maxSegmentSize, 
> maxOverlapSize and embeddingBatchSize do not apply; documentSplitter is 
> rejected
>   - id filters, dedup claim and release, documentFilter and IngestResult 
> outcomes work unchanged
>   Out of scope: chunking long audio into time windows, storing the MIME type 
> as metadata, sniffing the type from the bytes.
>   Note: no LangChain4j provider embeds audio out of the box yet, so the model 
> is typically a custom EmbeddingModel wrapping an audio embedding service. 
> Tests use a
>   deterministic audio-capable fake.
>   Follow-ups planned in camel-quarkus (extension option on camel-main) and 
> camel-kamelets (langchain4j-ingest-sink property).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to