Executive summary: Lots of false failures; many job's recheck'd; please be patient.
A bit more detail... Seems everybody has been very productive so far this week, submitting LOTS of diff's to review.opencontrail.org. Early last evening there were close to 80 reviews waiting for resources on the cluster, which was 100% utilized. Then bad things started to happen. The CI cluster ran out of 10.x floating-ip's several times, and that caused a lot of jobs to fail before they had even run. Many jobs hit this issue between 8pm and early this morning. Then github had an outage that started around 4:25am SVL time and continued for about 35min. This from github's status page: 11:32 UTC We're seeing high error rates on github.com and are investigating. 11:40 UTC We're doing emergency maintenance to recover the site. 11:54 UTC We've finished emergency maintenance and are monitoring closely. 12:20 UTC Everything operating normally. All jobs that started during this time quickly failed. I've recheck'd all the jobs impacted by this. Once again, the cluster is 100% busy, and there are >75 jobs queued and waiting. Ashish is trying to get us more resources for the cluster so that we can add more VM's. -- Klash _______________________________________________ Dev mailing list [email protected] http://lists.opencontrail.org/mailman/listinfo/dev_lists.opencontrail.org
