[ 
https://issues.apache.org/jira/browse/MESOS-3548?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15207331#comment-15207331
 ] 

Deepak Vij commented on MESOS-3548:
-----------------------------------

Nested/Hierarchical model typically employs something like an “Uber” Mesos 
Cluster as the control plane. This introduces a level of indirection which 
leads to various issues such as failure scenarios of actual Mesos clusters in 
the federation – the application failure information/state in Mesos Cluster 
level has to propagate back to Generic framework/s for appropriate action/s at 
the Uber controller level. Then there is complexity related to monitoring of 
deployed applications/services at Uber Mesos Cluster from the end user 
standpoint. As the Uber Mesos cluster is a regular Mesos cluster itself, what 
happens to the frameworks such as Marathon etc. which perform health-checks on 
the deployed applications etc. etc.

Instead, we decided to go with the multi-master approach whereby frameworks 
connect directly to multiple masters in the federation. Flow of offers to the 
frameworks in such a federated environment is controlled via replicated policy 
data store. This design is inspired by the Redis 3.0 Cluster’s design approach. 

We have had few iterations of our design reviews with some prominent folks in 
the Mesos community. This work is very strategic for our technology direction 
that allows us to support complex cloud computing use case scenarios such as 
“Cloud Bursting”, “Resilience”, avoid “Vendor Lock-in etc. Hope all this makes 
sense.

> Investigate federations of Mesos masters
> ----------------------------------------
>
>                 Key: MESOS-3548
>                 URL: https://issues.apache.org/jira/browse/MESOS-3548
>             Project: Mesos
>          Issue Type: Improvement
>            Reporter: Neil Conway
>              Labels: federation, mesosphere, multi-dc
>
> In a large Mesos installation, the operator might want to ensure that even if 
> the Mesos masters are inaccessible or failed, new tasks can still be 
> scheduled (across multiple different frameworks). HA masters are only a 
> partial solution here: the masters might still be inaccessible due to a 
> correlated failure (e.g., Zookeeper misconfiguration/human error).
> To support this, we could support the notion of "hierarchies" or 
> "federations" of Mesos masters. In a Mesos installation with 10k machines, 
> the operator might configure 10 Mesos masters (each of which might be HA) to 
> manage 1k machines each. Then an additional "meta-Master" would manage the 
> allocation of cluster resources to the 10 masters. Hence, the failure of any 
> individual master would impact 1k machines at most. The meta-master might not 
> have a lot of work to do: e.g., it might be limited to occasionally 
> reallocating cluster resources among the 10 masters, or ensuring that newly 
> added cluster resources are allocated among the masters as appropriate. 
> Hence, the failure of the meta-master would not prevent any of the individual 
> masters from scheduling new tasks. A single framework instance probably 
> wouldn't be able to use more resources than have been assigned to a single 
> Master, but that seems like a reasonable restriction.
> This feature might also be a good fit for a multi-datacenter deployment of 
> Mesos: each Mesos master instance would manage a single DC. Naturally, 
> reducing the traffic between frameworks and the meta-master would be 
> important for performance reasons in a configuration like this.
> Operationally, this might be simpler if Mesos processes were self-hosting 
> ([MESOS-3547]).



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Reply via email to