> I have already signed up on Apache Flink Confluence [1] under the > username Chanhae > Oh. > > If you are able to grant me write access, I will move the current draft > there using the FLIP template and proceed toward a community vote.
Done, you can have a try. Best, Leonard > > Best regards, > > Chanhae > > On Tue, Jul 21, 2026 at 12:16 PM Leonard Xu <[email protected]> wrote: > >> Hi ChanHae, >> >> Thanks for the update. The newly added limitations and boundary-query >> result >> requirements address my remaining concerns. No further concerns from me. >> >> One small procedural suggestion: FLIPs are usually maintained in the Apache >> Flink Confluence wiki, so it would be great if you could move the current >> Google Docs draft there using the FLIP template [1]. This will make it >> easier >> for the community to review, track, and eventually vote on the FLIP. >> >> If you do not have write permission to the Flink wiki yet, no worries. >> Please >> feel free to share your Apache Confluence username, and I can help add the >> permission for you. >> >> Best, >> Leonard >> >> [1] >> https://cwiki.apache.org/confluence/spaces/FLINK/pages/65145553/FLIP+Template >> >>> 2026 7月 20 11:41 下午,Chanhae Oh <[email protected]> 写道: >>> >>> Hi Leonard, >>> >>> Thank you for the +1 and for the additional suggestions. >>> >>> I have updated the FLIP based on your comments (changes marked with green >>> underlines): >>> >>> - Physical-id partitioning: added explicit limitation note for >>> static/write-free tables >>> - Boundary-query partitioning: added result requirements (single column, >>> type compatibility, NULL rejection) >>> >>> Please let me know if there is anything further needed before moving to >> the >>> next stage. >>> >>> Best regards, >>> Chanhae Oh >>> >>> On Mon, Jul 20, 2026 at 1:19 PM Leonard Xu <[email protected]> wrote: >>> >>>> Hi ChanHae, >>>> >>>> Thanks for the update. With these semantics clarified, I think the >>>> proposal is >>>> reasonable to move forward. >>>> >>>> My remaining suggestion is to make the limitations very explicit in both >>>> the >>>> FLIP and the user documentation. In particular, physical-id partitioning >>>> should >>>> be documented as suitable for static tables, or tables without >> concurrent >>>> changes during the scan. For boundary-query partitioning, the connector >>>> should >>>> reject NULL boundary values and require the boundary query to return a >>>> single >>>> column that is comparable with the partition column. >>>> >>>> With these documented, I am +1 on moving this FLIP forward. >>>> >>>> Best, >>>> Leonard >>>> >>>>> 2026 7月 8 7:12 下午,Chanhae Oh <[email protected]> 写道: >>>>> >>>>> Hi Leonard, >>>>> >>>>> Thank you for taking the time to review this proposal. >>>>> Your feedback touched on the exact boundary conditions and correctness >>>>> guarantees that I had not articulated precisely enough, and I genuinely >>>>> learned from the breadth of your perspective. >>>>> >>>>> I have updated the FLIP to address each of your questions (changed >>>> sections >>>>> are highlighted in yellow). >>>>> >>>>> 1. Physical ID Partitioning — correctness under concurrent changes >>>>> (Updated: Proposed Changes > Physical ID Partitioning > Correctness >>>>> Guarantees and Limitations) >>>>> >>>>> The two databases behave differently at the physical layer. Oracle >> ROWID >>>> is >>>>> stable under normal DML — it only changes under structural operations >>>>> (ALTER TABLE MOVE, EXPORT/IMPORT, row movement). PostgreSQL ctid, >>>> however, >>>>> changes on every UPDATE since a new tuple is written at a new location. >>>>> Physical-id partitioning for PostgreSQL is therefore only safe when no >>>>> concurrent DML runs during the scan. >>>>> >>>>> 2.1. Boundary-query — should the query return exactly N-1 values? >>>>> (Updated: Proposed Changes > Boundary Query > Partition count behavior) >>>>> >>>>> Requiring exactly N-1 is too strict. Fewer values are allowed — the >>>>> partition count simply reduces to (returned values) + 1. More than N-1 >>>>> values is a validation error, since scan.partition.num serves as a >>>> resource >>>>> cap and silently exceeding it could create more parallelism than the >> user >>>>> intended. >>>>> >>>>> 2.2. Boundary-query — how are inclusive/exclusive predicates generated? >>>>> (Updated: Proposed Changes > Boundary Query > Predicate contract) >>>>> >>>>> Partition 1 uses col < v1 OR col IS NULL, middle partitions use v(i-1) >> <= >>>>> col < vi (left-inclusive, right-exclusive), and the last partition uses >>>> col >>>>>> = v(N-1). This ensures no row is emitted by two partitions >>>> simultaneously. >>>>> >>>>> 2.3. Boundary-query — NULL values, duplicate boundaries, unsorted >>>> results, >>>>> count mismatch >>>>> (Updated: Proposed Changes > Boundary Query > Predicate contract, >>>> Partition >>>>> count behavior) >>>>> >>>>> - NULL: Routed to Partition 1 via the explicit col IS NULL clause, >> since >>>>> SQL NULL comparisons evaluate to NULL rather than TRUE. >>>>> - Duplicate boundary values: Handled gracefully by the predicate >> pattern >>>> — >>>>> the affected middle partition produces an empty result set with no data >>>>> loss. >>>>> - Unsorted results: The connector wraps the boundary-query in a >> subquery >>>>> with ORDER BY 1, so the user does not need to include ORDER BY. >>>>> - Count mismatch: Same as 2.1 — fewer values reduce partition count; >> more >>>>> values are rejected at validation. >>>>> >>>>> Additionally, I added an Option Validation Matrix section to make the >>>>> constraints across the three partitioning modes explicit in one place. >>>>> >>>>> Please let me know if any of the updated content raises further >>>> questions. >>>>> >>>>> Thank you again for your thorough review. >>>>> >>>>> P.S Thank you for your interest Youngil Kim. >>>>> >>>>> Best regards, >>>>> ChanHae Oh >>>>> >>>>> On Wed, Jul 8, 2026 at 2:48 PM Leonard Xu <[email protected]> wrote: >>>>> >>>>>> Hi Chanhae, >>>>>> >>>>>> Thanks for driving this FLIP. The motivation makes sense to me: the >>>> current >>>>>> equal-width JDBC partitioning does not work well for skewed data, and >> it >>>>>> also >>>>>> does not help much when users do not have a suitable >>>> numeric/date/timestamp >>>>>> partition column. >>>>>> >>>>>> I am generally supportive of the direction, but I think both proposed >>>>>> approaches >>>>>> need clearer semantics. >>>>>> >>>>>> For physical-id partitioning, ROWID/ctid are physical identifiers >> rather >>>>>> than >>>>>> stable logical keys. They may change after updates, table rewrites, >>>>>> vacuum, or >>>>>> row movement. Since different splits may be read by different >>>> statements or >>>>>> connections, the FLIP should clarify whether this mode only targets >>>> static >>>>>> tables, or whether the connector can provide a consistent snapshot >>>> across >>>>>> all >>>>>> splits. Otherwise, missing or duplicated rows may be possible. >>>>>> >>>>>> For boundary-query partitioning, the contract should also be more >>>> precise. >>>>>> For >>>>>> example, if scan.partition.num is N, should the query return N - 1 >>>>>> boundaries? >>>>>> How are inclusive/exclusive predicates generated? How should NULL >>>> values, >>>>>> duplicate boundaries, unordered results, or boundary count mismatches >> be >>>>>> handled? >>>>>> These details are important to avoid missing or duplicated rows. >>>>>> >>>>>> It would also be helpful to define the option validation matrix >> clearly, >>>>>> e.g. >>>>>> which options are required or mutually exclusive for range >> partitioning, >>>>>> boundary-query partitioning, and physical-id partitioning. >>>>>> >>>>>> Overall, I like the direction, but I think the FLIP should first nail >>>> down >>>>>> the >>>>>> correctness guarantees and edge-case behavior. >>>>>> >>>>>> Best, >>>>>> Leonard >>>> >>>> >> >>
