Amazon Corrects Massive AWS S3 Cloud Outage While Vendors React
Last Tuesday, parts of the Internet came to a grinding halt when the servers that powered them suddenly vanished. The disappearing server act came from servers that were housed as part of Amazon S3, Amazon’s popular Web hosting service.
When that incident happened, several big and popular services and Web sites were disrupted, including DraftKings, Gizmodo, IFTTT, Quora, Slack and Trello.
According to the Web site monitoring firm Apica, 54 of the largest online retailers experienced performance impairments on their Web sites, with some slowing down by more than 20 percent; 3 sites went down completely (Express, Lulu Lemon, One Kings Lane); and for effected websites, average slow down time was 29.7 seconds – 42.7 seconds to load.
What happened?
"At 9:37 a.m. PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process," Amazon said. "Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. The servers that were inadvertently removed supported two other S3 subsystems."
Those subsystems are important. One of them "manages the metadata and location information of all S3 objects in the region," according to Amazon. And without it, services that depend on it couldn’t perform basic data retrieval and storage tasks. The second subsystem, the placement subsystem, "manages allocation of new storage and requires the index subsystem to be functioning properly to correctly operate." The placement subsystem is used to allocate storage for new objects.



