Your website is down. How do you find the root cause?
From communication to log analysis, developers share their step-by step strategies.
• 4 min read
A website crash takes many forms. For example, WordPress Security Specialist Stefan Ristić recently helped a theater nonprofit deal with the slings and arrows of a 503 error (which indicate temporarily unavailable servers, often due to an overload of requests).
“For a few hours the site worked fine. Then, for a few hours, it did not work because the server could not serve those requests from the visitors. It was basically clogged with a ton of requests,” Ristić, also the founder of domain security research project TLDWP, told us.
Here are a few reasons why a site might experience an empty page, a spinning wheel, an error number, or a timeout:
- Server. Maybe it’s bogged down by bots, an unexpected traffic spike, or a hardware issue.
- Domain Name System (DNS). The DNS hosting service, which maps a domain to an IP address, may be disrupted; alternatively, a domain registration might have expired.
- Code. Even tiny coding errors can prevent important scripts from running; broken or deleted dependencies or software updates may likewise lead to outages.
- Certificates. An expired encryption cert like SSL or TLS may hit users with a security warning.
With so many possible causes for crashes, where does a site admin start when investigating a root cause?
An early step: check the logs. In Ristić’s case, he said he examined access logs and found about 1,000 requests in one minute from 997 IPs. (Bots continue to overwhelm sites and site visitors: cybersecurity company Imperva’s 2026 Bad Bot Report found that automated traffic accounted for more than 53% of all web traffic in 2025, up from 51% the year before.)
Ristić noticed the requests had targeted many variations of valid URL combinations. One of the website’s filtering plugins allowed multiple legitimate URL combinations to reach the site, based on date and category filters. The bots, perhaps trying to deploy a denial-of-service attack, hit large numbers of the URLs from many different IP addresses, making simple IP blocking impractical. Ristić found common characteristics with the URLs and set up a rule (via Cloudflare) that sent a managed challenge to suspicious requests for URLs of that category, blocking the traffic before reaching the server.
From cybersecurity and big data to cloud computing, IT Brew covers the latest trends shaping business tech in our 4x weekly newsletter, virtual events with industry experts, and digital guides.
By subscribing, you accept our Terms & Privacy Policy.
“The first step is always to check the error logs, because most of the time they will tell us exactly what’s going on and what we need to fix first,” Ristić said.
One tool being added to the mix these days is AI models. When investigating root crash causes, Illia Zub, VP of software engineering at web search API provider SerpApi, has lately asked AI tools (like Anthropic’s Opus 5.5 and OpenAI’s GPT-6 Astra) to review data like error logs, GitHub issues, recent code changes, and server metrics.
“I use AI to basically figure out the root cause, to isolate problems, and test different hypotheses,” Zub told us.
The steps. Senior Site Reliability Engineer Sai Joshitha Kathari spoke with IT Brew in August about how she handled a production incident early in her career. In a follow-up email in October, Kathari reviewed the steps that helped then, and are still relevant today in determining the cause of a disruption:
- Impact: Determine who’s affected, what services are down, and if the issue is getting worse.
- Get the data: Review dashboards, infrastructure signals, and platform events. Look at signals like application logs, recently deployment or configuration changes, and resource usage.
- Stabilize what you can safely: Take “clear and low-risk recovery action” where possible (like restarting a crashing pod or computing unit) while continuing the investigation.
- Communicate: Keep relevant engineers and app owners informed, and escalate when more expertise is required or the problem worsens.
- Verify the recovery: Look for errors, latency, and unexpected application behavior or instability.
- Document: Share the determined root cause, what fixed the problem, and what should change to prevent the issue from reoccurring.
“The biggest thing I learned wasn’t a technical skill at all,” Kathari told us in August, when reflecting on her early-career production problem. “It was how to stay calm.”
About the author
Billy Hurley
Billy Hurley has been a reporter with IT Brew since 2022. He writes stories about cybersecurity threats, AI developments, and IT strategies.
From cybersecurity and big data to cloud computing, IT Brew covers the latest trends shaping business tech in our 4x weekly newsletter, virtual events with industry experts, and digital guides.
By subscribing, you accept our Terms & Privacy Policy.