How to deal with cloud failure: Live, learn, fix, repeat
Grazed from GigaOM. Author: Derrick Harris.
Like it or not, sweeping software bugs are just part and parcel with operating the largest computing systems the world has ever seen. On Monday night, Amazon Web Services published a detailed post-mortem of its latest cloud outage, which struck on Friday night as massive thunderstorms knocked out power to one of the company’s east coast data centers. However, issues with the data center’s backup generator were just a catalyst — it was a handful of latent software bugs that manifested themselves as the system attempted to restore itself that did the real damage.
Although AWS is already working on fixes at all levels, this won’t be the last cloud computing outage we see, either from AWS or its competitors in the cloud provider space…

