International Edition
Latest News
Technology

It’s not just you: The internet is breaking

In the span of a few months this year, the internet has managed to knock itself sideways four different ways. And the official explanations have landed with all the romance of a maintenance log. A Cloudflare file exceeded…

It’s not just you: The internet is breaking

In the span of a few months this year, the internet has managed to knock itself sideways four different ways. And the official explanations have landed with all the romance of a maintenance log. A Cloudflare file exceeded its expected size. A DNS entry inside AWS pointed nowhere. An Azure configuration change went sideways. A Google service-control rule looped into failure and sent itself into repeated crash cycles.

each failure began as a routine maintenance task – the digital equivalents of leaving a door ajar. Each one expanded into a global interruption. 

These events slid into place quietly and revealed the same uncomfortable truth: The internet is a tightly bound structure, not a sprawling, distributed network, as many people may imagine.A small change in one corner sets off a chain reaction in another because so many digital services rely on the same gateways, the same load balancers, the same identity checkpoints, and the same routing layers. The fragility sits inside those shared pathways, not inside the individual apps that blinked out of view.

So,no,you’re not wrong: The internet feels like it’s breaking – because we’ve made it too big to fail and too small at the top to stay upright.## Tiny fixes become global problems

When a file exceeded its expected size at Cloudflare in November, it wasn’t a malicious attack. It was a configuration error. But because Cloudflare handles so much internet traffic – roughly 20% of all requests, according to the company – the error cascaded into outages for sites like Discord, Reddit, and Shopify.More than 17 million user-reported failures stacked up in the first hours.that number was large enough to show how dependent companies remain on AWS’ core regions – even when architects insist they have spread their risk. Region redundancy offered little insulation because identity checks, data calls, and background tasks still funnel through the most popular region by habit.The failure didn’t last long, but it still reached sectors that thought they stood outside the impact zone. Welcome to the modern cloud.

Azure’s turn arrived the following week when a traffic-management update in a Microsoft edge layer slowed down workplace logins, airline check-ins, retail portals, and gaming platforms. The surface symptoms looked disconnected. The underlying problem sat in a routing system tied to microsoft’s identity stack.Many organizations that don’t run their applications on Azure still rely on Microsoft to verify credentials,authorize sessions,or route user data. A shift in that layer appears small on paper. But in practice, it affects travel, commerce, communication, and office workflows – all at the same time.

A service-control rule slipped into the wrong layer inside Google Cloud over the summer and knocked the platform off balance.The code that signs off on routine API calls kept crashing and restarting, and requests that usually clear in a blink began to stall or fall away. The stutter showed up across regions as authentication failures, halted builds, and applications blinking in and out of view – hitting streaming platforms, collaboration tools, and Google’s own systems before the platform managed to steady itself. It didn’t last long, but it made plain that Google’s control plane behaves like a single surface, and a small shift in that layer follows every path that depends on it.

One web, one spine

These failures didn’t come from the same flaw. But they pointed to the same structure.“`html





The Growing Impact of cloud Outages


The Growing Impact of Cloud Outages

Recent incidents involving major cloud providers – AWS, Cloudflare, Azure, and Google Cloud – demonstrate a critical shift in how we assess the severity of outages. Analysts are increasingly focused not on how *long* an outage lasts, but on how *widely* it spreads.These events are no longer isolated incidents; they are capable of causing global disruptions affecting thousands of businesses and millions of users.

The Scale of recent Disruptions

The impact of these outages is significant. The AWS incident affected over 3,500 companies in more than 60 countries. cloudflare’s failure generated over 11,000 user-incident reports, impacting banks, retailers, logistics, media, and even government agencies. Azure’s slowdown saw over 30,000 outage reports within the first hour, disrupting travel, entertainment, and countless other digital services. Azure’s October 2025 outage was notably widespread. Google Cloud’s June 2025 incident further highlighted this trend.

Why Blast Radius Matters More

Traditionally, outage analysis centered on Mean Time to Recovery (MTTR). While minimizing downtime remains crucial, the sheer interconnectedness of modern digital infrastructure means a wider blast radius can have far more significant consequences. A short outage affecting a large number of critical services is far more damaging than a longer outage impacting a smaller user base.

  • Interdependence: Businesses increasingly rely on a complex web of cloud services. an outage in one service can cascade through multiple systems.
  • global Reach: Cloud services operate globally, meaning a single point of failure can impact users worldwide.
  • Critical Infrastructure: Many essential services – finance, healthcare, transportation – now depend on cloud infrastructure, making outages potentially life-altering.

Understanding the Root Causes

These outages aren’t simply random occurrences. Several factors contribute to their increasing frequency and impact:

Complexity of Cloud Systems

Cloud environments are incredibly complex, involving vast amounts of code, intricate configurations, and numerous interconnected components.This complexity makes it difficult to identify and resolve issues quickly.

Software Bugs and Configuration Errors

Many outages are caused by software bugs or misconfigurations. Even minor errors can have widespread consequences in a large-scale cloud surroundings.

Increased Demand and Scalability Challenges

Rapid growth in cloud adoption puts strain on infrastructure, making it harder to maintain stability and scalability. unexpected surges in demand can overwhelm systems and led to outages.

Mitigating the Risk

While eliminating outages entirely is unrealistic, organizations can take steps to mitigate the risk and minimize the impact:

  • Multi-Cloud Strategy: Distributing workloads across multiple cloud providers reduces reliance on any single vendor.
  • Redundancy and Failover: Implementing redundant systems and automated failover mechanisms ensures business continuity in the event of an outage.
  • Robust Monitoring and Alerting: Proactive monitoring and alerting systems can detect issues early and enable faster response times.
  • Disaster recovery Planning: Having a well-defined disaster recovery plan is essential for restoring services quickly after an outage.

Key Takeaways

  • The focus is shifting from outage *duration* to outage *blast radius
About the author: Anika Shah - Technology

MSc in Computer Science, senior reporter. Anika focuses on AI ethics, cybersecurity, and emerging hardware—frequently moderating panels at CES and Web Summit. “Anika Shah decodes tech breakthroughs and startup disruption shaping tomorrow’s digital landscape.”