[ https://issues.apache.org/jira/browse/SPARK-1405?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14303918#comment-14303918 ]
Joseph K. Bradley edited comment on SPARK-1405 at 2/3/15 8:29 PM: ------------------------------------------------------------------ Thanks everyone for all of your contributions, help and feedback! The initial LDA implementation has been merged, but there are many improvements which remain to be done. I've put a list of JIRAs here [https://issues.apache.org/jira/browse/SPARK-5572] [~yuhaoyan] +1 for online LDA. (I made a JIRA for it.) was (Author: josephkb): Thanks everyone for all of your contributions, help and feedback! The initial LDA implementation has been merged, but there are many improvements which remain to be done. I've put a list of JIRAs here [https://issues.apache.org/jira/browse/SPARK-5572] [~yuhao yang] +1 for online LDA. (I made a JIRA for it.) > parallel Latent Dirichlet Allocation (LDA) atop of spark in MLlib > ----------------------------------------------------------------- > > Key: SPARK-1405 > URL: https://issues.apache.org/jira/browse/SPARK-1405 > Project: Spark > Issue Type: New Feature > Components: MLlib > Reporter: Xusen Yin > Assignee: Joseph K. Bradley > Priority: Critical > Labels: features > Fix For: 1.3.0 > > Attachments: performance_comparison.png > > Original Estimate: 336h > Remaining Estimate: 336h > > Latent Dirichlet Allocation (a.k.a. LDA) is a topic model which extracts > topics from text corpus. Different with current machine learning algorithms > in MLlib, instead of using optimization algorithms such as gradient desent, > LDA uses expectation algorithms such as Gibbs sampling. > In this PR, I prepare a LDA implementation based on Gibbs sampling, with a > wholeTextFiles API (solved yet), a word segmentation (import from Lucene), > and a Gibbs sampling core. -- This message was sent by Atlassian JIRA (v6.3.4#6332) --------------------------------------------------------------------- To unsubscribe, e-mail: issues-unsubscr...@spark.apache.org For additional commands, e-mail: issues-h...@spark.apache.org