[
https://issues.apache.org/jira/browse/HIVE-16552?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15992308#comment-15992308
]
Rui Li commented on HIVE-16552:
-------------------------------
[~xuefuz], my point is a large number of tasks doesn't mean a lot of resources.
For example, a user can request only 1 container with 1 slot, and submit a job
containing 1000 tasks. At any point there'll be no more than 1 task running
simultaneously. On the other hand, a user can also request lots of containers
but only run small jobs with them. The latter user of course takes up more
resources and has bigger impact on other users.
> Limit the number of tasks a Spark job may contain
> -------------------------------------------------
>
> Key: HIVE-16552
> URL: https://issues.apache.org/jira/browse/HIVE-16552
> Project: Hive
> Issue Type: Improvement
> Components: Spark
> Affects Versions: 1.0.0, 2.0.0
> Reporter: Xuefu Zhang
> Assignee: Xuefu Zhang
> Attachments: HIVE-16552.1.patch, HIVE-16552.patch
>
>
> It's commonly desirable to block bad and big queries that takes a lot of YARN
> resources. One approach, similar to mapreduce.job.max.map in MapReduce, is to
> stop a query that invokes a Spark job that contains too many tasks. The
> proposal here is to introduce hive.spark.job.max.tasks with a default value
> of -1 (no limit), which an admin can set to block queries that trigger too
> many spark tasks.
> Please note that this control knob applies to a spark job, though it's
> possible that one query can trigger multiple Spark jobs (such as in case of
> map-join). Nevertheless, the proposed approach is still helpful.
--
This message was sent by Atlassian JIRA
(v6.3.15#6346)