Routing system failure

AWS outage knocks thousands of websites offline worldwide

AWS
Facebook
X
LinkedIn
Reddit
WhatsApp
image source: JHVEPhoto/Shutterstock.com

A routing failure in AWS CloudFront caused global outages, impacting thousands of services across North America, Europe, and Asia for over three hours.

Amazon Web Services has resolved a massive, three-and-a-half-hour outage of its CloudFront content delivery network. The disruption began at 12:45 a.m. Pacific Daylight Time (PDT) on Thursday and resulted in thousands of websites around the globe becoming inaccessible and returning 5xx errors. The issue exclusively affected customers using Virtual Private Cloud Origins (VPC Origins) to connect public websites with private cloud resources.

Ad

According to system reports, the cause lay within a packet processing subsystem that was unable to forward requests from CloudFront edge locations to resources inside the customers’ private networks. An internal constraint on managing connections to private VPC Origins was reached. Consequently, the system responsible for distributing routing configurations to network processors failed to load the updated configuration data, completely blocking traffic.

https://twitter.com/piyokango/status/2077690631280058440?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E2077690631280058440%7Ctwgr%5E0ce9be45dc395fbe05826ac68dac32b5e1e0906a%7Ctwcon%5Es1_&ref_url=https%3A%2F%2Fcybernews.com%2Fnews%2Faws-cloudfront-outage-websites-5xx-errors%2F

Global impact on businesses and services

Because the underlying VPC technology is deployed globally, the disruption spanned across all continents. According to market data, more than 6,000 companies in North America alone use the affected Amazon services. In Europe, there are more than 2,100, and in Asia, over 1,300 companies are affected.

Concretely confirmed outages included major platforms and services in various regions. In Europe, the UK National Lottery was among those affected. Additionally, the AI developer portal HuggingFace, the networking service Tailscale, and the hardware and software manufacturer Ubiquiti reported severe accessibility issues and error messages with their cloud services. Following the phased deployment of an update by Amazon engineers, all systems were restored to normal operations.

Ad
https://twitter.com/AWSSupport/status/2077678987443077449?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E2077678987443077449%7Ctwgr%5E0ce9be45dc395fbe05826ac68dac32b5e1e0906a%7Ctwcon%5Es1_&ref_url=https%3A%2F%2Fcybernews.com%2Fnews%2Faws-cloudfront-outage-websites-5xx-errors%2F

Criticism of communication policy and cloud centralization

In the wake of the global incident, industry experts highlighted the risks associated with the growing centralization of the cloud market. Mayur Upadhyaya, CEO of APIContext, commented on the structural issues of this consolidation:

“The AWS CloudFront incident isn’t noteworthy because a critical cloud provider experienced an outage. Every infrastructure provider will have operational incidents.”

Mayur Upadhyaya, CEO of APIContext

He added that a single system failure now impacts thousands of businesses simultaneously because so many business models depend on the exact same infrastructure. In addition to technical vulnerability, numerous users on social media expressed frustration over AWS’s crisis communication. The company failed to update its status dashboard during the initial hours of the outage, leaving administrators of the affected platforms in the dark regarding the cause of the 5xx errors.

(red)

Ad

Weitere Artikel