+1

Best,
Hongbo



原始邮件
发件人:JUNHAO YE <[email protected]>
发件时间:2026年7月7日 17:12
收件人:[email protected] <[email protected]>
主题:Re: [DISCUSS] Apache Paimon 2.0.0 release


+1

Best,
Junhao

> 2026年7月7日 15:50,wj wang <[email protected]> 写道:
> 
> +1 for the proposal. Looking forward to Paimon 2.0!
> 
> Best,
> wangwj
> 
> On Tue, Jul 7, 2026 at 3:24 PM Yunfeng Zhou <[email protected]> 
>wrote:
>> 
>> Hi Jingsong,
>> 
>> +1 for the proposal. Looking forward to Paimon 2.0!
>> 
>> Best,
>> Yunfeng
>> 
>>> 2026年7月7日 11:11,Jingsong Li <[email protected]> 写道:
>>> 
>>> Hi everyone,
>>> 
>>> I would like to start a discussion about preparing an Apache Paimon
>>> 2.0.0 release.
>>> 
>>> Over the past release cycles, Paimon has evolved significantly. In
>>> addition to continuing improvements to the core lakehouse table
>>> engine, we have been building a broader storage foundation for
>>> streaming, analytics, and AI/multimodal workloads.
>>> 
>>> Some major areas that have landed or are being stabilized include:
>>> 
>>> 1. Data Evolution and row tracking
>>> 
>>> Paimon now has a much stronger foundation for append tables with row
>>> tracking and data evolution. This enables efficient partial column
>>> updates, schema evolution, dedicated storage for special column types,
>>> and global-index-based retrieval without rewriting entire data files.
>>> 
>>> 2. Multimodal table capabilities
>>> 
>>> We have introduced multimodal table support, covering blob storage,
>>> vector storage, full-text content, and global indexes in one table
>>> abstraction. This allows Paimon tables to store and query structured
>>> data together with images, videos, audio, embeddings, and text
>>> content.
>>> 
>>> 3. Global Index framework
>>> 
>>> The global index framework has been extended to support multiple index
>>> types, including BTree, Bitmap, vector indexes, full-text indexes, and
>>> hybrid search. This provides a unified mechanism for scalar filtering,
>>> vector similarity search, text retrieval, and combined retrieval
>>> workflows.
>>> 
>>> 4. Vector and full-text search
>>> 
>>> We have added native vector index support, including paimon-vindex
>>> IVF-based indexes and Lumina DiskANN-based indexes, as well as native
>>> full-text search integration with BM25 scoring and tokenizer
>>> configuration. These features are important for RAG, recommendation,
>>> image retrieval, and other AI-native scenarios.
>>> 
>>> 5. New storage formats and supporting components
>>> 
>>> We have also made several supporting components available around
>>> Paimon, including Mosaic, paimon-vindex, and paimon-full-text-index.
>>> 
>>> Mosaic provides a columnar-bucket hybrid file format optimized for
>>> wide tables and efficient column projection. paimon-vindex provides
>>> native vector indexing capabilities for Paimon. paimon-full-text-index
>>> provides the native full-text engine used by Paimon full-text global
>>> indexes.
>>> 
>>> Together, these components make Paimon more capable as a storage layer
>>> for both traditional lakehouse workloads and emerging AI/multimodal
>>> workloads.
>>> 
>>> 6. Ecosystem improvements
>>> 
>>> We have continued to improve Flink, Spark, Python, Ray, and Daft
>>> integration, and have been adding more end-to-end coverage for
>>> Java/Python interoperability, global index creation and reading,
>>> vector data, full-text search, and data evolution workflows.
>>> 
>>> Given the scope of these changes, I think it is a good time to discuss
>>> whether we should cut a 2.0.0 release. A 2.0 release would help
>>> communicate that Paimon has entered a new stage, with data evolution,
>>> multimodal storage, and global indexes becoming first-class parts of
>>> the project.
>>> 
>>> Please share your thoughts, concerns, or any issues that you believe
>>> should block or be included in the 2.0.0 release.
>>> 
>>> Best,
>>> Jingsong
>> 


Reply via email to