[
https://issues.apache.org/jira/browse/FLINK-40323?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Thorne updated FLINK-40323:
---------------------------
Description:
*1、Background & Problem Description*
In real-time data synchronization scenarios, most business databases widely use
*JSON/JSONB data types* to store flexible, unstructured and semi-structured
business data, such as user extension attributes, order additional information,
device metadata and dynamic business parameters.
Currently, Flink CDC connectors lack native and standardized support for
database JSON data types. For table fields defined as JSON/JSONB in source
databases, the existing processing logic has the following problems: *
JSON fields are forcibly parsed as plain string types, losing original JSON
structural attribute information
*
Inconsistent JSON parsing rules across different database connectors, resulting
in data synchronization deviation
*
Cannot dock with Flink-CDC native JSON data type, failing to support subsequent
JSON function calculations and structured analysis
* Why not use variant data type to replace , because some sink connectors do
not support variant data type and i think that we should let the pipeline job
more precise.
The above limitations greatly restrict the capability of Flink CDC in
semi-structured data real-time synchronization and analysis, and cannot meet
the flexible business data synchronization demands of modern data warehouses
and real-time computing platforms.
h3. 2、 Requirement
We need to add *native JSON data type support* for Flink CDC core connectors,
to implement full-process standardized parsing, identification and
synchronization of database JSON/JSONB fields, and perfectly dock with Flink
built-in JSON data types and JSON computing functions.
h3. 3、Core Capabilities
* {*}Native Type Identification{*}: Automatically identify JSON/JSONB type
fields in source database tables during schema parsing, distinguish them from
ordinary string fields
* {*}Standardized Data Parsing{*}: Parse original database JSON binary/string
data into Flink native JSON type objects, retain complete JSON structure
(key-value pairs, nested structures, arrays)
* {*}Cross-database Compatibility{*}: Uniform parsing rules for JSON types of
mainstream databases (MySQL JSON, PostgreSQL JSONB, etc.) to eliminate
cross-database parsing differences
* {*}Complete Change Capture{*}: Support capturing full JSON field changes
(full update of JSON content, partial update of nested fields) in CDC
incremental synchronization
* {*}Compatible with Existing Logic{*}: Provide compatible configuration
items, support fallback to string parsing mode, ensure no impact on existing
online job
h2. 4. Related Tasks
*
Optimize CDC schema parsing logic to add JSON type mapping rules
*
Rewrite JSON field deserialization logic for each connector
*
Add compatible configuration parameters and documentation updates
*
Supplement end-to-end test cases for JSON data type synchronizatio
was:
*1、Background & Problem Description*
In real-time data synchronization scenarios, most business databases widely use
*JSON/JSONB data types* to store flexible, unstructured and semi-structured
business data, such as user extension attributes, order additional information,
device metadata and dynamic business parameters.
Currently, Flink CDC connectors lack native and standardized support for
database JSON data types. For table fields defined as JSON/JSONB in source
databases, the existing processing logic has the following problems: *
JSON fields are forcibly parsed as plain string types, losing original JSON
structural attribute information
*
Inconsistent JSON parsing rules across different database connectors, resulting
in data synchronization deviation
*
Cannot dock with Flink-CDC native JSON data type, failing to support subsequent
JSON function calculations and structured analysis
The above limitations greatly restrict the capability of Flink CDC in
semi-structured data real-time synchronization and analysis, and cannot meet
the flexible business data synchronization demands of modern data warehouses
and real-time computing platforms.
h3. 2、 Requirement
We need to add *native JSON data type support* for Flink CDC core connectors,
to implement full-process standardized parsing, identification and
synchronization of database JSON/JSONB fields, and perfectly dock with Flink
built-in JSON data types and JSON computing functions.
h3. 3、Core Capabilities
*
{*}Native Type Identification{*}: Automatically identify JSON/JSONB type fields
in source database tables during schema parsing, distinguish them from ordinary
string fields
*
{*}Standardized Data Parsing{*}: Parse original database JSON binary/string
data into Flink native JSON type objects, retain complete JSON structure
(key-value pairs, nested structures, arrays)
*
{*}Cross-database Compatibility{*}: Uniform parsing rules for JSON types of
mainstream databases (MySQL JSON, PostgreSQL JSONB, etc.) to eliminate
cross-database parsing differences
*
{*}Complete Change Capture{*}: Support capturing full JSON field changes (full
update of JSON content, partial update of nested fields) in CDC incremental
synchronization
*
{*}Compatible with Existing Logic{*}: Provide compatible configuration items,
support fallback to string parsing mode, ensure no impact on existing online job
h2. 4. Related Tasks
*
Optimize CDC schema parsing logic to add JSON type mapping rules
*
Rewrite JSON field deserialization logic for each connector
*
Add compatible configuration parameters and documentation updates
*
Supplement end-to-end test cases for JSON data type synchronizatio
> [cdc-common] Support JSON Data Type in Flink CDC
> ------------------------------------------------
>
> Key: FLINK-40323
> URL: https://issues.apache.org/jira/browse/FLINK-40323
> Project: Flink
> Issue Type: Improvement
> Components: Flink CDC
> Reporter: Thorne
> Priority: Major
>
> *1、Background & Problem Description*
> In real-time data synchronization scenarios, most business databases widely
> use *JSON/JSONB data types* to store flexible, unstructured and
> semi-structured business data, such as user extension attributes, order
> additional information, device metadata and dynamic business parameters.
> Currently, Flink CDC connectors lack native and standardized support for
> database JSON data types. For table fields defined as JSON/JSONB in source
> databases, the existing processing logic has the following problems: *
> JSON fields are forcibly parsed as plain string types, losing original JSON
> structural attribute information
> *
> Inconsistent JSON parsing rules across different database connectors,
> resulting in data synchronization deviation
> *
> Cannot dock with Flink-CDC native JSON data type, failing to support
> subsequent JSON function calculations and structured analysis
> * Why not use variant data type to replace , because some sink connectors do
> not support variant data type and i think that we should let the pipeline job
> more precise.
> The above limitations greatly restrict the capability of Flink CDC in
> semi-structured data real-time synchronization and analysis, and cannot meet
> the flexible business data synchronization demands of modern data warehouses
> and real-time computing platforms.
> h3. 2、 Requirement
> We need to add *native JSON data type support* for Flink CDC core connectors,
> to implement full-process standardized parsing, identification and
> synchronization of database JSON/JSONB fields, and perfectly dock with Flink
> built-in JSON data types and JSON computing functions.
> h3. 3、Core Capabilities
> * {*}Native Type Identification{*}: Automatically identify JSON/JSONB type
> fields in source database tables during schema parsing, distinguish them from
> ordinary string fields
> * {*}Standardized Data Parsing{*}: Parse original database JSON
> binary/string data into Flink native JSON type objects, retain complete JSON
> structure (key-value pairs, nested structures, arrays)
> * {*}Cross-database Compatibility{*}: Uniform parsing rules for JSON types
> of mainstream databases (MySQL JSON, PostgreSQL JSONB, etc.) to eliminate
> cross-database parsing differences
> * {*}Complete Change Capture{*}: Support capturing full JSON field changes
> (full update of JSON content, partial update of nested fields) in CDC
> incremental synchronization
> * {*}Compatible with Existing Logic{*}: Provide compatible configuration
> items, support fallback to string parsing mode, ensure no impact on existing
> online job
>
> h2. 4. Related Tasks
> *
> Optimize CDC schema parsing logic to add JSON type mapping rules
> *
> Rewrite JSON field deserialization logic for each connector
> *
> Add compatible configuration parameters and documentation updates
> *
> Supplement end-to-end test cases for JSON data type synchronizatio
--
This message was sent by Atlassian Jira
(v8.20.10#820010)