Hi Jingsong,

+1 for the proposal. Looking forward to Paimon 2.0!

Best,
Yunfeng

> 2026年7月7日 11:11,Jingsong Li <[email protected]> 写道:
> 
> Hi everyone,
> 
> I would like to start a discussion about preparing an Apache Paimon
> 2.0.0 release.
> 
> Over the past release cycles, Paimon has evolved significantly. In
> addition to continuing improvements to the core lakehouse table
> engine, we have been building a broader storage foundation for
> streaming, analytics, and AI/multimodal workloads.
> 
> Some major areas that have landed or are being stabilized include:
> 
> 1. Data Evolution and row tracking
> 
> Paimon now has a much stronger foundation for append tables with row
> tracking and data evolution. This enables efficient partial column
> updates, schema evolution, dedicated storage for special column types,
> and global-index-based retrieval without rewriting entire data files.
> 
> 2. Multimodal table capabilities
> 
> We have introduced multimodal table support, covering blob storage,
> vector storage, full-text content, and global indexes in one table
> abstraction. This allows Paimon tables to store and query structured
> data together with images, videos, audio, embeddings, and text
> content.
> 
> 3. Global Index framework
> 
> The global index framework has been extended to support multiple index
> types, including BTree, Bitmap, vector indexes, full-text indexes, and
> hybrid search. This provides a unified mechanism for scalar filtering,
> vector similarity search, text retrieval, and combined retrieval
> workflows.
> 
> 4. Vector and full-text search
> 
> We have added native vector index support, including paimon-vindex
> IVF-based indexes and Lumina DiskANN-based indexes, as well as native
> full-text search integration with BM25 scoring and tokenizer
> configuration. These features are important for RAG, recommendation,
> image retrieval, and other AI-native scenarios.
> 
> 5. New storage formats and supporting components
> 
> We have also made several supporting components available around
> Paimon, including Mosaic, paimon-vindex, and paimon-full-text-index.
> 
> Mosaic provides a columnar-bucket hybrid file format optimized for
> wide tables and efficient column projection. paimon-vindex provides
> native vector indexing capabilities for Paimon. paimon-full-text-index
> provides the native full-text engine used by Paimon full-text global
> indexes.
> 
> Together, these components make Paimon more capable as a storage layer
> for both traditional lakehouse workloads and emerging AI/multimodal
> workloads.
> 
> 6. Ecosystem improvements
> 
> We have continued to improve Flink, Spark, Python, Ray, and Daft
> integration, and have been adding more end-to-end coverage for
> Java/Python interoperability, global index creation and reading,
> vector data, full-text search, and data evolution workflows.
> 
> Given the scope of these changes, I think it is a good time to discuss
> whether we should cut a 2.0.0 release. A 2.0 release would help
> communicate that Paimon has entered a new stage, with data evolution,
> multimodal storage, and global indexes becoming first-class parts of
> the project.
> 
> Please share your thoughts, concerns, or any issues that you believe
> should block or be included in the 2.0.0 release.
> 
> Best,
> Jingsong

Reply via email to