Every so often, the digital world gets a reminder that “The Cloud” is not a magical, omnipresent entity but just another complex socio-technical system operated by humans. Electricity usually comes from the socket, networks usually work, and cloud services are marketed as being always on… Until they are not.
The recent large-scale Cloudflare outage was one of those reminders. A significant portion of internet-facing services suddenly became unreachable, not because of a cyberattack or a natural disaster, but because of an internal failure in a widely used platform. For many organizations, the most unsettling part was not even the outage itself, but the realization that there was nothing they could do. No fallback button, no emergency procedure, just waiting for status updates.
Outages with global impact have been happening with uncomfortable frequency recently. AWS and Azure have both demonstrated how a failure in a single region or service can cascade into disruptions far beyond what customers expected. Even Cloudflare itself experienced another, shorter outage only weeks after the major incident. Individually, these events may look like bad days. Taken together, they point to a structural problem.
From a Resilient Internet to Fragile Overlays
The original internet was designed to survive catastrophic failures. It was never meant to be perfectly reliable, but it was meant to be resilient. Traffic could take different paths, components could fail without collapsing the entire system, and there was no single control plane that everything depended on.
What we have built on top of that foundation looks very different. Content delivery networks, cloud platforms, security services, identity providers, and API gateways act as powerful overlays that concentrate enormous amounts of traffic and functionality. They improve performance, simplify operations, and enable business models that would otherwise be impossible. At the same time, they reintroduce single points of failure, just at a much larger scale.
It is also worth dispelling one persistent myth. There is simply no single global cloud, just as there is no single global internet. Clouds are proprietary ecosystems operated by individual providers, segmented into regions, governed by local regulations, and increasingly shaped by geopolitical realities. Networks can and do become disconnected. Countries can and do impose controls. Treating “the cloud” as a homogeneous, always-on utility is an oversimplification that leads to dangerous assumptions.
Platform Dependence and the Risk We Keep Ignoring
The appeal of platforms is obvious. One provider with a single contract and centralized control plane. Security, networking, performance optimization, and analytics integrated into a coherent whole. For many organizations, this reduces cost and complexity in the short term.
The trade-off is dependency. The more functionality is consolidated into a single platform, the harder it becomes to operate without it, even temporarily. When a platform fails, customers often discover that their carefully designed redundancy exists only inside that platform. DNS (somehow, it’s always DNS!), security controls, identity services, and management interfaces all disappear at once.
What makes this particularly problematic is that outages of this magnitude are no longer exceptional. Yet many architectural and procurement decisions still implicitly assume that such failures are either extremely unlikely or somebody else’s problem. Risk management tends to focus on efficiency, performance, and cost, while systemic failure scenarios are quietly deprioritized.
Rethinking Resilience: Some Practical Conclusions
The lesson from recent cloud outages is not that platforms or hyperscalers are inherently flawed. They deliver real value and are unlikely to disappear. The lesson is that resilience cannot be outsourced entirely and that systemic risk needs to be addressed more explicitly. A few conclusions seem increasingly hard to ignore.
First, large cloud platforms need to invest more seriously in internal resilience. This includes better isolation between services, more rigorous change management, staged rollouts with effective kill switches, and architectural designs that prevent a single configuration error from propagating globally. Transparency around incidents and their root causes should be mandated as a core responsibility.
Second, enterprises must reassess how much dependency on a single provider or platform they are willing to accept. Multi-cloud strategies, selective use of best-of-breed services, and avoiding unnecessary platform lock-in all come at a cost, but so does being unable to operate for hours because one external service failed. Exit strategies, fallback options, and realistic failure scenarios should be part of architectural decisions.
Finally, there is a broader question for regulators and end users. When outages at a handful of providers can disrupt large parts of the economy, healthcare, or public services, it is reasonable to ask whether they should be treated as critical infrastructure. That may imply stronger regulatory oversight, resilience requirements, and accountability mechanisms. Pushing this discussion forward is not anti-cloud. It is a recognition of how central these services have become.
Digital Sovereignty and the End of the Global Internet Illusion
There is one additional dimension that makes this discussion even more complex: digital sovereignty. An increasing number of countries are pursuing some form of digital independence, driven by legitimate concerns around data protection, national security, and economic autonomy. The practical outcome is new legislation and regulatory pressure that limits the role of global, typically US-based, service providers in favor of local alternatives.
From a policy perspective, this may make sense, but it introduces new architectural trade-offs. Local cloud providers and national platforms often operate at a much smaller scale, with limited geographic distribution and fewer redundancy options. While they may offer advantages in terms of jurisdictional control and compliance, they are typically inherently less resilient than hyperscalers that operate dozens of regions worldwide.
In September 2025, South Korea experienced a major loss of government data following a fire at the National Information Resources Service data center. This was not a failure of a global hyperscaler but of a nationally critical, sovereign infrastructure component. The incident demonstrated that sovereignty does not automatically translate into resilience and that concentrating critical services in a smaller ecosystem can increase, not reduce systemic risk.
Fragmentation as Risk and as Opportunity
These developments raise uncomfortable questions. If the Internet and the Cloud are becoming increasingly fragmented, does this further undermine the idea of a globally resilient digital infrastructure? Will customers face even more outages as they are forced to rely on smaller, more isolated ecosystems? And does multi-cloud still help if those clouds are no longer globally interoperable?
The pessimistic answer is that fragmentation will indeed lead to more complexity and more failure scenarios. The more optimistic view is that it may finally force organizations to confront architectural reality. Resilience has always required deliberate design, explicit trade-offs, and acceptance that failure is normal. Digital sovereignty initiatives make it harder to outsource these decisions and easier to see the consequences of ignoring them.
For enterprises, this can be an opportunity to reassess priorities. Instead of optimizing solely for convenience or regulatory compliance, they may need to balance sovereignty, resilience, cost, and operational complexity more consciously. For providers, both global and local, it should be a signal that availability and failure containment are as important as feature breadth and market reach.
The Cloud is not going away, and neither are outages. The global Internet as a single, uniform entity may quickly become more myth than reality. The real question is whether we continue to design systems the old way, or whether we finally start building architectures that acknowledge fragmentation, failure, and responsibility as first-class design constraints.
For those looking to explore these questions in more depth, events like EIC 2026 in Berlin next May provide an opportunity to exchange experiences and discuss emerging best practices with practitioners and experts. In a world of recurring outages and growing fragmentation, that kind of shared learning may be one of the most effective resilience measures we still have.