Vinoth #1 sounds amazing, I'm having a bit of a time visualizing it, would you be able to provide some concrete examples?
On Mon, Mar 4, 2019 at 10:02 PM Vinoth Chandar <[email protected]> wrote: > Hello all, > > Bubbling up and summarizing a very fascinating discussion we had a while > back with Nick Vintila around leveraging Hudi to incrementalize aspects of > ML pipelines.. > > Specifically, > > 1. Currently ML feature engineering is done via a combination of batch and > streaming pipelines. If we can distill some common patterns on which > features are windowed/aggregated, we could support "merge hooks" inside > Hudi for them and in turn compute a lot of batch features incrementally. > Apache Beam integration is a good first step here with Hudi as a IO, to > provide incremental change streams & absorb updates. > > 2. Second large aspect is providing a way to "safely" join different hudi > datasets to compute these features. In other words, stream/stream > stream/table joins are hard and to get consistent/correct results, one > needs to know completeness information for these tables and trigger joins > only when such criteria is met. On this front, should Hudi provide multi > dataset transactionality - is one school of thought. > > I will let Nick chip in more with probably better clarity :) > > Thanks > Vinoth >
