Hi Leonard,
Thank you for confirming — and for facilitating the Confluence access so smoothly. The FLIP[1] wiki page is now in place (Yuepeng Pan helped set it up, and I have completed the edits). I will be moving forward with the vote shortly. [1] cwiki.apache.org/confluence/x/x4Q_Gg Best regards, ChanHae Oh 2026년 7월 22일 수요일, Leonard Xu <[email protected]>님이 작성: > > > I have already signed up on Apache Flink Confluence [1] under the > > username Chanhae > > Oh. > > > > If you are able to grant me write access, I will move the current draft > > there using the FLIP template and proceed toward a community vote. > > Done, you can have a try. > > Best, > Leonard > > > > > > Best regards, > > > > Chanhae > > > > On Tue, Jul 21, 2026 at 12:16 PM Leonard Xu <[email protected]> wrote: > > > >> Hi ChanHae, > >> > >> Thanks for the update. The newly added limitations and boundary-query > >> result > >> requirements address my remaining concerns. No further concerns from me. > >> > >> One small procedural suggestion: FLIPs are usually maintained in the > Apache > >> Flink Confluence wiki, so it would be great if you could move the > current > >> Google Docs draft there using the FLIP template [1]. This will make it > >> easier > >> for the community to review, track, and eventually vote on the FLIP. > >> > >> If you do not have write permission to the Flink wiki yet, no worries. > >> Please > >> feel free to share your Apache Confluence username, and I can help add > the > >> permission for you. > >> > >> Best, > >> Leonard > >> > >> [1] > >> https://cwiki.apache.org/confluence/spaces/FLINK/pages/ > 65145553/FLIP+Template > >> > >>> 2026 7月 20 11:41 下午,Chanhae Oh <[email protected]> 写道: > >>> > >>> Hi Leonard, > >>> > >>> Thank you for the +1 and for the additional suggestions. > >>> > >>> I have updated the FLIP based on your comments (changes marked with > green > >>> underlines): > >>> > >>> - Physical-id partitioning: added explicit limitation note for > >>> static/write-free tables > >>> - Boundary-query partitioning: added result requirements (single > column, > >>> type compatibility, NULL rejection) > >>> > >>> Please let me know if there is anything further needed before moving to > >> the > >>> next stage. > >>> > >>> Best regards, > >>> Chanhae Oh > >>> > >>> On Mon, Jul 20, 2026 at 1:19 PM Leonard Xu <[email protected]> wrote: > >>> > >>>> Hi ChanHae, > >>>> > >>>> Thanks for the update. With these semantics clarified, I think the > >>>> proposal is > >>>> reasonable to move forward. > >>>> > >>>> My remaining suggestion is to make the limitations very explicit in > both > >>>> the > >>>> FLIP and the user documentation. In particular, physical-id > partitioning > >>>> should > >>>> be documented as suitable for static tables, or tables without > >> concurrent > >>>> changes during the scan. For boundary-query partitioning, the > connector > >>>> should > >>>> reject NULL boundary values and require the boundary query to return a > >>>> single > >>>> column that is comparable with the partition column. > >>>> > >>>> With these documented, I am +1 on moving this FLIP forward. > >>>> > >>>> Best, > >>>> Leonard > >>>> > >>>>> 2026 7月 8 7:12 下午,Chanhae Oh <[email protected]> 写道: > >>>>> > >>>>> Hi Leonard, > >>>>> > >>>>> Thank you for taking the time to review this proposal. > >>>>> Your feedback touched on the exact boundary conditions and > correctness > >>>>> guarantees that I had not articulated precisely enough, and I > genuinely > >>>>> learned from the breadth of your perspective. > >>>>> > >>>>> I have updated the FLIP to address each of your questions (changed > >>>> sections > >>>>> are highlighted in yellow). > >>>>> > >>>>> 1. Physical ID Partitioning — correctness under concurrent changes > >>>>> (Updated: Proposed Changes > Physical ID Partitioning > Correctness > >>>>> Guarantees and Limitations) > >>>>> > >>>>> The two databases behave differently at the physical layer. Oracle > >> ROWID > >>>> is > >>>>> stable under normal DML — it only changes under structural operations > >>>>> (ALTER TABLE MOVE, EXPORT/IMPORT, row movement). PostgreSQL ctid, > >>>> however, > >>>>> changes on every UPDATE since a new tuple is written at a new > location. > >>>>> Physical-id partitioning for PostgreSQL is therefore only safe when > no > >>>>> concurrent DML runs during the scan. > >>>>> > >>>>> 2.1. Boundary-query — should the query return exactly N-1 values? > >>>>> (Updated: Proposed Changes > Boundary Query > Partition count > behavior) > >>>>> > >>>>> Requiring exactly N-1 is too strict. Fewer values are allowed — the > >>>>> partition count simply reduces to (returned values) + 1. More than > N-1 > >>>>> values is a validation error, since scan.partition.num serves as a > >>>> resource > >>>>> cap and silently exceeding it could create more parallelism than the > >> user > >>>>> intended. > >>>>> > >>>>> 2.2. Boundary-query — how are inclusive/exclusive predicates > generated? > >>>>> (Updated: Proposed Changes > Boundary Query > Predicate contract) > >>>>> > >>>>> Partition 1 uses col < v1 OR col IS NULL, middle partitions use > v(i-1) > >> <= > >>>>> col < vi (left-inclusive, right-exclusive), and the last partition > uses > >>>> col > >>>>>> = v(N-1). This ensures no row is emitted by two partitions > >>>> simultaneously. > >>>>> > >>>>> 2.3. Boundary-query — NULL values, duplicate boundaries, unsorted > >>>> results, > >>>>> count mismatch > >>>>> (Updated: Proposed Changes > Boundary Query > Predicate contract, > >>>> Partition > >>>>> count behavior) > >>>>> > >>>>> - NULL: Routed to Partition 1 via the explicit col IS NULL clause, > >> since > >>>>> SQL NULL comparisons evaluate to NULL rather than TRUE. > >>>>> - Duplicate boundary values: Handled gracefully by the predicate > >> pattern > >>>> — > >>>>> the affected middle partition produces an empty result set with no > data > >>>>> loss. > >>>>> - Unsorted results: The connector wraps the boundary-query in a > >> subquery > >>>>> with ORDER BY 1, so the user does not need to include ORDER BY. > >>>>> - Count mismatch: Same as 2.1 — fewer values reduce partition count; > >> more > >>>>> values are rejected at validation. > >>>>> > >>>>> Additionally, I added an Option Validation Matrix section to make the > >>>>> constraints across the three partitioning modes explicit in one > place. > >>>>> > >>>>> Please let me know if any of the updated content raises further > >>>> questions. > >>>>> > >>>>> Thank you again for your thorough review. > >>>>> > >>>>> P.S Thank you for your interest Youngil Kim. > >>>>> > >>>>> Best regards, > >>>>> ChanHae Oh > >>>>> > >>>>> On Wed, Jul 8, 2026 at 2:48 PM Leonard Xu <[email protected]> wrote: > >>>>> > >>>>>> Hi Chanhae, > >>>>>> > >>>>>> Thanks for driving this FLIP. The motivation makes sense to me: the > >>>> current > >>>>>> equal-width JDBC partitioning does not work well for skewed data, > and > >> it > >>>>>> also > >>>>>> does not help much when users do not have a suitable > >>>> numeric/date/timestamp > >>>>>> partition column. > >>>>>> > >>>>>> I am generally supportive of the direction, but I think both > proposed > >>>>>> approaches > >>>>>> need clearer semantics. > >>>>>> > >>>>>> For physical-id partitioning, ROWID/ctid are physical identifiers > >> rather > >>>>>> than > >>>>>> stable logical keys. They may change after updates, table rewrites, > >>>>>> vacuum, or > >>>>>> row movement. Since different splits may be read by different > >>>> statements or > >>>>>> connections, the FLIP should clarify whether this mode only targets > >>>> static > >>>>>> tables, or whether the connector can provide a consistent snapshot > >>>> across > >>>>>> all > >>>>>> splits. Otherwise, missing or duplicated rows may be possible. > >>>>>> > >>>>>> For boundary-query partitioning, the contract should also be more > >>>> precise. > >>>>>> For > >>>>>> example, if scan.partition.num is N, should the query return N - 1 > >>>>>> boundaries? > >>>>>> How are inclusive/exclusive predicates generated? How should NULL > >>>> values, > >>>>>> duplicate boundaries, unordered results, or boundary count > mismatches > >> be > >>>>>> handled? > >>>>>> These details are important to avoid missing or duplicated rows. > >>>>>> > >>>>>> It would also be helpful to define the option validation matrix > >> clearly, > >>>>>> e.g. > >>>>>> which options are required or mutually exclusive for range > >> partitioning, > >>>>>> boundary-query partitioning, and physical-id partitioning. > >>>>>> > >>>>>> Overall, I like the direction, but I think the FLIP should first > nail > >>>> down > >>>>>> the > >>>>>> correctness guarantees and edge-case behavior. > >>>>>> > >>>>>> Best, > >>>>>> Leonard > >>>> > >>>> > >> > >> > >
