commits
Thread
Date
Earlier messages
Later messages
Messages by Thread
Re: [PR] docs(tech-specs): update the delete block format for Avro serde [hudi]
via GitHub
[I] hudi-azure package dependencies affecting hudi write operations [hudi]
via GitHub
Re: [I] [SUPPORT] Data loss on V5 MOR-Avro table with Flink async instant generation: compaction plan races with inflight deltacommit [hudi]
via GitHub
(hudi) branch master updated: docs(hudi-io): fill in the HFile format details the doc was missing (#19721)
vhs
Re: [I] Update hfile_format.md accordingly [hudi]
via GitHub
Re: [PR] fix(clean): use completion time for KEEP_LATEST_BY_HOURS retention [hudi]
via GitHub
Re: [PR] fix(clean): use completion time for KEEP_LATEST_BY_HOURS retention [hudi]
via GitHub
(hudi) branch master updated: fix(hive-sync): pass the default partition through the slash-encoded value extractors (#19710)
danny0405
Re: [I] Support Flink 2.2 [hudi]
via GitHub
Re: [I] [BUG] SlashEncodedDayPartitionValueExtractor throws on __HIVE_DEFAULT_PARTITION__, breaking hive sync of slash tables with null partition values [hudi]
via GitHub
[PR] feat(core): read HFile base files through ranged, window-coalesced reads [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through ranged, window-coalesced reads [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through ranged, window-coalesced reads [hudi-rs]
via GitHub
Re: [I] All Azure CI jobs should use same linux version in the test runner image [hudi]
via GitHub
Re: [I] All Azure CI jobs should use same linux version in the test runner image [hudi]
via GitHub
Re: [I] Support INSERT SQL with a subset of columns and select from in Spark 3.5 [hudi]
via GitHub
Re: [I] Support INSERT SQL with a subset of columns and select from in Spark 3.5 [hudi]
via GitHub
Re: [I] Improve get file size for writer [hudi]
via GitHub
Re: [I] Improve get file size for writer [hudi]
via GitHub
Re: [I] Support data skipping for flink streaming source [hudi]
via GitHub
Re: [I] Support data skipping for flink streaming source [hudi]
via GitHub
Re: [I] Investigate Azure CI flakey test custom detection [hudi]
via GitHub
Re: [I] Investigate Azure CI flakey test custom detection [hudi]
via GitHub
Re: [I] Close stable PRs [hudi]
via GitHub
Re: [I] Close stable PRs [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [PR] fix(spark-sql): lay meta fields over the partial-update schema in the global-index merge [hudi]
via GitHub
Re: [I] Auto-infer variant shredding schema (SPARK + AVRO record types) [hudi]
via GitHub
Re: [I] Add auto shredding field inference logic [hudi]
via GitHub
Re: [I] Fix Trino integration tests in docker demo [hudi]
via GitHub
Re: [I] Fix Trino integration tests in docker demo [hudi]
via GitHub
(hudi) branch master updated: fix(spark-sql): resolve a partition path without validating the record key (#19709)
vhs
[PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] perf(flink): preempt inactive write buckets on memory exhaustion [hudi]
via GitHub
Re: [PR] fix(utilities): use endOffsets when no offset is greater than the checkpoint timestamp [hudi]
via GitHub
Re: [PR] fix(utilities): use endOffsets when no offset is greater than the checkpoint timestamp [hudi]
via GitHub
(hudi) branch bot/trino-pin updated (7e6c79203f3e -> a4db03a5a5bb)
github-bot
(hudi) 01/01: chore(trino): advance trino master pin to 8fbf2737d4c4
github-bot
Re: [PR] fix(spark): make partition DDL commands honor slash separated date partitioning [hudi]
via GitHub
Re: [PR] fix(spark): make partition DDL commands honor slash separated date partitioning [hudi]
via GitHub
Re: [PR] fix(spark): make partition DDL commands honor slash separated date partitioning [hudi]
via GitHub
Re: [PR] fix(spark): make partition DDL commands honor slash separated date partitioning [hudi]
via GitHub
Re: [PR] fix(spark): make partition DDL commands honor slash separated date partitioning [hudi]
via GitHub
Re: [PR] fix(spark): make partition DDL commands honor slash separated date partitioning [hudi]
via GitHub
Re: [PR] fix(common): rejoin slash-separated partition values for non-time-typed columns [hudi]
via GitHub
[PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
Re: [PR] feat(spark): support bucket index for LSM tables [hudi]
via GitHub
(hudi) branch master updated: fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager (#19718)
danny0405
(hudi) branch master updated (8adb095386e3 -> 81fd4992b8dc)
danny0405
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
Re: [I] [Discussion WIP] Add utility to both schedule and execute table services on MDT without needing to write to data table [hudi]
via GitHub
[PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
Re: [PR] test(sync): cover the meta-sync completion-time watermark end to end [hudi]
via GitHub
[I] Meta-sync completion-time watermark has no test coverage and the Glue test fixture commit is invisible to the timeline [hudi]
via GitHub
Re: [I] Meta-sync completion-time watermark has no test coverage and the Glue test fixture commit is invisible to the timeline [hudi]
via GitHub
Re: [PR] chore(api): declare the unstructured ingestion SPIs evolving [hudi]
via GitHub
Re: [I] Unstructured ingestion SPIs are unannotated, so their stability is undeclared [hudi]
via GitHub
(hudi) branch master updated: chore(api): declare the unstructured ingestion SPIs evolving (#19701)
yihua
Re: [PR] docs: RFC-110: propose Hudi full-text search index [hudi]
via GitHub
Re: [PR] docs: RFC-110: propose Hudi full-text search index [hudi]
via GitHub
[PR] feat(core): read HFile base files through the version two reader [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through the version two reader [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through the version two reader [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through the version two reader [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through the version two reader [hudi-rs]
via GitHub
Re: [PR] feat(core): read HFile base files through the version two reader [hudi-rs]
via GitHub
[PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
Re: [PR] fix(spark): skip record key config validation for table version 1 [hudi]
via GitHub
[I] Table version 1 tables cannot be written to on the 0.x line: record key config conflict blocks the upgrade that would fix it [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(spark-sql): resolve MERGE INTO partition columns so records are not mis-partitioned [hudi]
via GitHub
Re: [PR] fix(hive-sync): pass the default partition through the slash-encoded value extractors [hudi]
via GitHub
Re: [PR] fix(hive-sync): pass the default partition through the slash-encoded value extractors [hudi]
via GitHub
Re: [PR] fix(utilities): ship Tika's parsers in the bundle and fail when they are absent [hudi]
via GitHub
Re: [PR] fix(utilities): harden the unstructured file source and embedding transformer [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(hive-sync): call Driver.destroy() so HiveQL sync stops leaking Drivers into ShutdownHookManager [hudi]
via GitHub
Re: [PR] fix(core): reject a second INSERT_OVERWRITE of a partition already overwritten by a concurrent writer [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [PR] fix(metadata): indexing action initializes the partition it asked for [hudi]
via GitHub
Re: [I] Slash-separated date partitioning throws ClassCastException on the Spark row writer path [hudi]
via GitHub
(hudi) branch master updated: fix(spark): make slash-separated date partitioning work on the row writer path (#19648)
vhs
Re: [PR] [HUDI-9624] Warn when event-time field is set without watermark tracking [hudi]
via GitHub
Re: [PR] [HUDI-9624] Warn when event-time field is set without watermark tracking [hudi]
via GitHub
Re: [PR] [HUDI-9624] Warn when event-time field is set without watermark tracking [hudi]
via GitHub
Re: [PR] [HUDI-9624] Warn when event-time field is set without watermark tracking [hudi]
via GitHub
Re: [PR] [HUDI-9624] Warn when event-time field is set without watermark tracking [hudi]
via GitHub
Re: [PR] [HUDI-9624] Warn when event-time field is set without watermark tracking [hudi]
via GitHub
Re: [PR] fix(variant): close the shredded-read gaps exposed by a mixed-layout test matrix [hudi]
via GitHub
Re: [PR] fix(variant): close the shredded-read gaps exposed by a mixed-layout test matrix [hudi]
via GitHub
Earlier messages
Later messages