[
https://issues.apache.org/jira/browse/CAMEL-25027?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jiri Ondrusek updated CAMEL-25027:
----------------------------------
Summary: camel-langchain4j-ingest - Add media modality to embed images,
audio, video and PDF whole (was: camel-langchain4j-ingest - Add audio
modality to embed audio documents whole)
> camel-langchain4j-ingest - Add media modality to embed images, audio, video
> and PDF whole
> --------------------------------------------------------------------------------------------
>
> Key: CAMEL-25027
> URL: https://issues.apache.org/jira/browse/CAMEL-25027
> Project: Camel
> Issue Type: New Feature
> Components: camel-langchain4j
> Affects Versions: 4.23.0
> Reporter: Jiri Ondrusek
> Assignee: Jiri Ondrusek
> Priority: Major
>
> The langchain4j-ingest producer only handles text: the body is read as a
> String, split into segments and embedded with embedAll. Audio files cannot be
> ingested as audio
> embeddings, the way a Wav2Vec2 or CLAP model would embed them for acoustic
> similarity search. Without a parser they are decoded as garbage text and
> embedded anyway.
> LangChain4j 1.19 has the API for this: EmbeddingRequest accepts
> AudioContent and a model declares AUDIO in supportedContentTypes().
>
> Proposal: a modality endpoint option, text (default, unchanged behaviour)
> or audio.
> With modality=audio:
> - the body is read as bytes and embedded whole, as one vector, through
> embed(EmbeddingRequest) with an AudioContent
> - the vector is stored with a placeholder TextSegment whose text is the
> document id, carrying the same camel_ingest_pipeline and
> camel_ingest_document_id metadata as text
> segments
> - the endpoint fails to start when the model does not declare AUDIO
> - the MIME type comes from a new contentType option or, when unset, from
> the document id's file extension (wav, mp3, flac, ogg, m4a, aac); an unknown
> type fails the
> exchange
> - maxDocumentSize and minDocumentSize count bytes; maxSegmentSize,
> maxOverlapSize and embeddingBatchSize do not apply; documentSplitter is
> rejected
> - id filters, dedup claim and release, documentFilter and IngestResult
> outcomes work unchanged
> Out of scope: chunking long audio into time windows, storing the MIME type
> as metadata, sniffing the type from the bytes.
> Note: no LangChain4j provider embeds audio out of the box yet, so the model
> is typically a custom EmbeddingModel wrapping an audio embedding service.
> Tests use a
> deterministic audio-capable fake.
> Follow-ups planned in camel-quarkus (extension option on camel-main) and
> camel-kamelets (langchain4j-ingest-sink property).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)