+1 Best, Junhao
> 2026年7月7日 15:50,wj wang <[email protected]> 写道: > > +1 for the proposal. Looking forward to Paimon 2.0! > > Best, > wangwj > > On Tue, Jul 7, 2026 at 3:24 PM Yunfeng Zhou <[email protected]> > wrote: >> >> Hi Jingsong, >> >> +1 for the proposal. Looking forward to Paimon 2.0! >> >> Best, >> Yunfeng >> >>> 2026年7月7日 11:11,Jingsong Li <[email protected]> 写道: >>> >>> Hi everyone, >>> >>> I would like to start a discussion about preparing an Apache Paimon >>> 2.0.0 release. >>> >>> Over the past release cycles, Paimon has evolved significantly. In >>> addition to continuing improvements to the core lakehouse table >>> engine, we have been building a broader storage foundation for >>> streaming, analytics, and AI/multimodal workloads. >>> >>> Some major areas that have landed or are being stabilized include: >>> >>> 1. Data Evolution and row tracking >>> >>> Paimon now has a much stronger foundation for append tables with row >>> tracking and data evolution. This enables efficient partial column >>> updates, schema evolution, dedicated storage for special column types, >>> and global-index-based retrieval without rewriting entire data files. >>> >>> 2. Multimodal table capabilities >>> >>> We have introduced multimodal table support, covering blob storage, >>> vector storage, full-text content, and global indexes in one table >>> abstraction. This allows Paimon tables to store and query structured >>> data together with images, videos, audio, embeddings, and text >>> content. >>> >>> 3. Global Index framework >>> >>> The global index framework has been extended to support multiple index >>> types, including BTree, Bitmap, vector indexes, full-text indexes, and >>> hybrid search. This provides a unified mechanism for scalar filtering, >>> vector similarity search, text retrieval, and combined retrieval >>> workflows. >>> >>> 4. Vector and full-text search >>> >>> We have added native vector index support, including paimon-vindex >>> IVF-based indexes and Lumina DiskANN-based indexes, as well as native >>> full-text search integration with BM25 scoring and tokenizer >>> configuration. These features are important for RAG, recommendation, >>> image retrieval, and other AI-native scenarios. >>> >>> 5. New storage formats and supporting components >>> >>> We have also made several supporting components available around >>> Paimon, including Mosaic, paimon-vindex, and paimon-full-text-index. >>> >>> Mosaic provides a columnar-bucket hybrid file format optimized for >>> wide tables and efficient column projection. paimon-vindex provides >>> native vector indexing capabilities for Paimon. paimon-full-text-index >>> provides the native full-text engine used by Paimon full-text global >>> indexes. >>> >>> Together, these components make Paimon more capable as a storage layer >>> for both traditional lakehouse workloads and emerging AI/multimodal >>> workloads. >>> >>> 6. Ecosystem improvements >>> >>> We have continued to improve Flink, Spark, Python, Ray, and Daft >>> integration, and have been adding more end-to-end coverage for >>> Java/Python interoperability, global index creation and reading, >>> vector data, full-text search, and data evolution workflows. >>> >>> Given the scope of these changes, I think it is a good time to discuss >>> whether we should cut a 2.0.0 release. A 2.0 release would help >>> communicate that Paimon has entered a new stage, with data evolution, >>> multimodal storage, and global indexes becoming first-class parts of >>> the project. >>> >>> Please share your thoughts, concerns, or any issues that you believe >>> should block or be included in the 2.0.0 release. >>> >>> Best, >>> Jingsong >>
