[ 
https://issues.apache.org/jira/browse/SPARK-6830?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14606176#comment-14606176
 ] 

Perinkulam I Ganesh commented on SPARK-6830:
--------------------------------------------

This thought crossed our mind as well earlier. So we were debating whether the 
caching should be implemented within the cacheManager, so that the count is 
cached only if the underlying RDD is cached.

> Memoize frequently queried vals in RDD, such as numPartitions, count etc.
> -------------------------------------------------------------------------
>
>                 Key: SPARK-6830
>                 URL: https://issues.apache.org/jira/browse/SPARK-6830
>             Project: Spark
>          Issue Type: Improvement
>          Components: SparkR
>            Reporter: Shivaram Venkataraman
>            Priority: Minor
>              Labels: Starter
>
> We should memoize frequently queried vals in RDD, such as numPartitions, 
> count etc.
> While using SparkR in RStudio, the `count` function seems to be called 
> frequently by the IDE – I think this is to show some stats about variables in 
> the workspace etc. but this is not great in SparkR as we trigger a job every 
> time count is called.
> Memoization would help in this case, but we should also see if there is some 
> better way to interact with RStudio.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

---------------------------------------------------------------------
To unsubscribe, e-mail: issues-unsubscr...@spark.apache.org
For additional commands, e-mail: issues-h...@spark.apache.org

Reply via email to