Technology

What Actually Happens When a Website Goes Down?

Or: why "the website is down" doesn't actually tell you what broke. A modern site is a chain of systems, and any single link failing can make the whole thing look dead.

Or: Why "The Website Is Down" Doesn't Actually Tell You What Broke

You click a bookmark. Nothing happens. You refresh. Still nothing. Maybe you get a blank page, an endlessly spinning browser tab, or one of those wonderfully reassuring messages: Server Not Found. 502 Bad Gateway. 503 Service Unavailable. Something went wrong.

You check another website. It works fine. So you reach the obvious conclusion: their server is down.

Maybe. But a modern website isn't one server sitting in a closet somewhere. It's a chain — DNS providers, content delivery networks, load balancers, web servers, application servers, databases, caches, authentication services, third-party APIs, payment processors — all cooperating to produce the page you see. Break the right link in that chain and the whole thing can appear dead, even while almost everything else is running fine.

"Down" Isn't a Single State

Before asking what broke, it helps to know what still works. Maybe the site fails on your home Wi-Fi but loads fine over cellular. Maybe the homepage loads but logging in fails. Maybe the text appears but every image is missing. Maybe checkout is broken while browsing works perfectly. All of these get reported as "it's down" — but they describe very different failures, and figuring out which one you're looking at is the first real step in troubleshooting.

The Problem Could Be You

This isn't an insult — your own device and network are part of the chain too. A failed Wi-Fi connection, a broken DNS resolver, a VPN routed through infrastructure having its own outage, a misbehaving browser extension, a stale cache, or an overzealous firewall can all make a perfectly healthy website look broken from where you're sitting. That's why the simplest test is opening a different website. If nothing loads anywhere, twenty unrelated companies didn't lose their servers simultaneously — your side of the connection just became the interesting part of the story. Trying another device, and then another network, narrows things further: change one variable at a time until you find the line between "works" and "doesn't."

When the Address Book Is Wrong

Before your browser can even reach a website, it typically needs DNS to translate the domain name into something usable. If a site's web servers, database, and application are all perfectly healthy but its DNS records are missing or misconfigured, the average visitor still sees a dead site. The building is open. The address book is wrong.

This gets stranger because DNS information is cached throughout the internet, so different users can genuinely have different experiences at the same moment. If a company changes its DNS records and something goes wrong, some resolvers may already have picked up the bad information while others are still serving the older, working answer from cache. One person says "it's down," another says "works fine for me" — and they're both telling the truth, just from different vantage points in a system that takes time to catch up with itself. That spreading-out process is often called DNS propagation, and it's one reason a DNS change can produce oddly regional or intermittent failures for a while.

Even once DNS resolves correctly, your traffic still has to physically reach the destination across networks run by your ISP and others. A fiber cut, a failed router, a bad peering relationship — any of these can make a perfectly healthy server unreachable from part of the internet while it's completely fine from elsewhere. The internet's resilience comes from having multiple possible paths between two points, but rerouting around a failure isn't always instant or perfect.

The CDN Layer Adds Its Own Failure Modes

Many websites sit behind a content delivery network, meaning your browser often isn't talking directly to the site's own servers at all. If that CDN has a problem, the website can look unavailable even though its origin servers are perfectly healthy — and because a single CDN provider often serves thousands of unrelated sites, an outage there can make a news site, an online store, and a completely unrelated game service all appear to break at the same time. The sites aren't connected. Their infrastructure is.

These failures can also be regional: a single edge location within a CDN can have trouble while users routed elsewhere never notice anything, producing that maddening back-and-forth of "it's down!" followed immediately by "no it isn't." And sometimes the CDN itself is fine but can't reach the origin behind it — which is exactly what a 502 Bad Gateway or 504 Gateway Timeout is trying to tell you. According to the MDN status code reference, a 502 generally means a server acting as a gateway received an invalid response from an upstream system, while a 504 means that upstream system simply never answered in time. The server showing you the error may be working perfectly; it's just reporting that a conversation behind it failed. A 500 Internal Server Error, by contrast, usually means the server itself hit an unexpected problem handling the request, and a 503 means the service is temporarily unable to handle traffic — often overload or maintenance. To a visitor, all four just mean "the page doesn't work." To someone debugging it, they point at very different layers.

Redundancy Is the Whole Point

Large services rarely run on one machine, because one machine is a single point of failure waiting to happen. Instead, traffic gets spread across many servers by a load balancer, which — per AWS's own description of the concept — distributes incoming requests across multiple targets based on their health and capacity, so no single server gets overwhelmed while others sit idle. Load balancers typically run continuous health checks, quietly pulling any server that stops responding correctly out of rotation so it can be repaired without taking the whole site down. This is why large services can absorb individual hardware failures constantly without users ever noticing — the system was built expecting them.

Of course, that raises the obvious question: what happens if the load balancer itself fails? The honest answer is that important infrastructure needs to be redundant too — multiple load balancers, DNS-based traffic distribution, managed cloud routing, entire backup data centers. The deeper you go, the more reliable computing turns out to be the art of repeatedly asking "what happens when this thing breaks?" and making sure the answer is never "everything dies."

Deployments Are a Dangerous Moment

Websites change constantly — bug fixes, new features, dependency updates — and every change is an opportunity for something to go wrong. Sometimes "the site went down" really means a team shipped something at 2:03 p.m. and immediately regretted it. That's why safer deployment patterns exist. In a canary deployment, a new version rolls out to a small slice of traffic first — maybe 1%, then 10%, then everyone — so a problem shows up as a warning sign rather than a full outage. In a blue-green deployment, two complete environments run side by side; traffic gets switched to the new one once it's verified, and switched right back if something's wrong. The underlying principle in both cases is the same: don't tear down the working bridge before you're sure the new one can carry traffic.

When the Slow Thing Breaks Everything Else

One of the stranger properties of large systems is that "slow" and "down" eventually become the same thing. Imagine a database that normally answers in 20 milliseconds suddenly taking 20 seconds. Nothing has technically crashed — but requests pile up, connection pools fill, and impatient users start refreshing, which generates more requests, which grows the backlog further. This is how a cascading failure happens: Service A depends on Service B, which depends on a database that's slowed down, and the slowdown ripples outward until several previously healthy systems are struggling because of one bottleneck none of them directly caused.

Automatic retries, meant to help, can make this dramatically worse. If a million users' applications all retry a failing request at once, the already-struggling server now faces a second wave larger than the first — a retry storm. Well-built systems fight this with timeouts (so nothing waits forever for a dependency that isn't answering), randomized backoff, and increasingly, circuit breakers that temporarily stop calling a failing dependency altogether, giving it room to recover instead of drowning it in retries.

Scale Doesn't Eliminate Complexity — It Multiplies It

Sometimes nothing is actually broken; too many people simply showed up. A site built for ten thousand visitors a minute can be overwhelmed by a million, even with every individual server technically functioning — connections fill, queues grow, requests start timing out. Modern infrastructure can often respond with autoscaling, adding more servers as demand rises, but scaling isn't instant and isn't even: adding fifty web servers doesn't help if all fifty are waiting on one overloaded database. A system is generally only as strong as its narrowest bottleneck. That same traffic pattern — an enormous, sudden surge — is also what a deliberate DDoS attack looks like from the infrastructure's point of view; the difference is intent, not symptoms, and either way the servers have to survive it.

Availability Has a Price Tag

Reliability gets described in "nines" — the percentage of time a service stays up — and each additional nine is dramatically harder to earn than the last. Google's own site reliability engineering documentation lays out the math plainly: 99% availability still allows for roughly 3.65 days of downtime across a year, while 99.9% brings that down to about 8.76 hours, 99.99% to about 52.6 minutes, and the famous "five nines," 99.999%, to roughly 5.26 minutes a year. Reaching each additional nine generally means another layer of redundancy — more regions, more automated failover, more real-time synchronization — and each layer costs real money and adds real complexity. Zero downtime isn't really the goal, because it isn't realistic. The actual goal is that individual pieces can fail without the whole service failing with them.

Recovery Isn't Instant Either

Even once engineers find and fix the actual problem, the system often doesn't snap back to normal immediately. Thousands of queued requests can flood a just-restarted database all at once. Caches that emptied during the outage need to refill, meaning requests that were normally served in microseconds temporarily have to hit slower systems again — a cold cache problem that can leave a service technically back online well before it's actually fast again. When something fails outright, failover can shift traffic to a healthy replica automatically, but even that has to be built carefully: two systems simultaneously believing they're the primary is its own kind of disaster.

People Are Part of the System Too

Hardware fails, software has bugs, and humans occasionally type the wrong command into the wrong environment. Some of the largest outages in internet history started with an ordinary mistake — a configuration pushed globally, a permission changed incorrectly, an engineer working in production while believing it was a test environment. Automation reduces a lot of that risk by applying consistent changes instead of manual ones — but automation applies a wrong configuration just as efficiently as a correct one, which is why mature systems pair automation with testing, staged rollouts, and rollback plans rather than trusting it blindly.

Good incident response also depends on visibility. Companies rely on monitoring that checks not just "is the server powered on" but whether the things users actually need — logging in, loading a page, completing a purchase — genuinely work; testing that only pings a machine's pulse can miss a completely dead database sitting right behind it. Public status pages, like the ones companies build with tools such as Atlassian's Statuspage, exist to communicate what's actually happening during an incident — ideally hosted independently from the infrastructure they're reporting on, to avoid the special irony of the status page going down along with the site it's supposed to describe.

The Bard's Take

When someone says "the website is down," all that sentence really confirms is that something between the user and the thing they wanted stopped working. It could be their own Wi-Fi. It could be a DNS resolver serving stale information. It could be a CDN edge location having a bad day, a load balancer misrouting traffic, an application that crashed, a database buckling under a slow query, a third-party payment API that's unreachable, an expired certificate, or one slow component that quietly dragged three healthy systems down with it.

That's why reliability engineering was never really about building computers that never fail — computers fail, drives fail, networks drop packets, and people make mistakes. The actual discipline is building systems where failure doesn't automatically become catastrophe: redundancy where it matters, traffic spread across many machines, sensible timeouts, limited retries, tested failover, gradual deployments, and constant monitoring of what users actually experience rather than just whether a machine is powered on.

The internet makes a website feel like one simple thing. You type an address, a page appears, and behind that page dozens of systems operated by several different companies across multiple data centers quietly do their part fast enough that you never notice any of them. Until one doesn't. Then the browser spins, you hit refresh, and you announce the universal diagnosis of the modern age: the website is down. The interesting question is always which part of that invisible chain made the sentence true.

Sources