Skip to content

Why Does a Website Suddenly Go Down on a Server?

When the site goes down, check DNS, network, TLS, Nginx, application and database layer by layer, save evidence, and recover the service with the least risk.

Author Bipida Editorial Team Published
Share this article

When the site suddenly goes down, the urgency to restart all services is understandable; but it also increases the scope of the failure and eliminates the evidence of cause. The definite may be DNS, network, TLS, reverse proxy, program, database, full disk, or memory shortage.

Quick answer:Record the start time and the exact user message, measure the DNS and TCP/TLS/HTTP status paths from multiple points and stop the last change. Then go inside the DNS, network, proxy, program, and dependency from outside. If recovery is necessary, save a snapshot of the log and status of the resources before that and make a change at each step.

First safety and communication, then flaw.

Specify the incident commander or decision maker so that no one changes the config at the same time. Record the times, commands, results and decisions in the timeline. If it is public, give a short, realistic status message; do not announce the time of the removal or the confirmed cause. Do not post access and secrecy on the public channel.

Down for everyone or just one user?

  • Test from a mobile network and an independent route.
  • Compare the main domain and a direct controlled hostname.
  • See IPv4 and IPv6 separately if they're active.
  • Separate the logged in user and the anonymous user.
  • Check all pages or just the dynamic operation.

The error of an ISP, browser cache, or local DNS differs from a public domain. On the other hand, the integrity of the cached homepage does not prove Checkout or API is healthy.

Convert the error message to layer

The sign.Possible layersFirst check.
The domain will not be resolvedDNS, registrar, DNSSEC, what is it?A valid record and name server.
Connection refusedListener, firewall, servicePort from outside andss
Connection timeoutRoutes, firewalls, providers.Network path and provider status
TLS errorCertificate, SNI, chain, timeHostname and validity of certificate
502/504Proxy and upstream.Error log and backend status
500Planning and Dependenceapplication/PHP log

A low-risk snapshot from the server.

date --iso-8601=seconds
uptime
free -h
df -h
df -i
ss -lntp
systemctl --failed

This output shows time, load, memory, disk, inode, listener and failed units and does not change. Delete IP information, private hostname, and confidential service name before subscribing. A snapshot of the moment should be compared to a full-time monitoring.

Step one: DNS

Check the A/AAAA record, CNAME, nameserver and DNSSEC status from standalone resolvers. Recent changes, expired records, or IPv6 references to unprepared servers can only disconnect a portion of users. Low TTL will speed up the release but will not fix the record error. Changing multiple records simultaneously makes it difficult to detect.

Phase two: Network and Port.

See that the IP is responsive and that the 80/443 ports are externally accessible. Ping is not a definite HTTP standard and may be intentionally closed. If there is a connection on the listener but it is not connected externally, check the cloud firewall, host firewall, route or provider. Do not change the management rule without an emergency console.

Step three: TLS

A certificate must be provided for the correct hostname, in a validity range, and with a full chain. The wrong clock will also spoil the validity of the server or client. If the renewal is done but the service provides an old certificate, the path or reload may be incorrect. HSTS can make temporary bypass HTTP impossible for the user; do not manipulate it without the program.

Step four: Nginx or web server.

Check the service status, config test, listener and error log. Active process does not mean the correct route; make a local request with the correct host. If the new configuration has an invalid syntax, the reload may have been rejected and the old process may still be active, or may not rise again after restarting.Configuring Nginx and WordPress.Look at that.

Step five: schedule and runtime

PHP-FPM, application server or container must be running and ready to respond. See log startup, exit code, restart count and healthcheck.502 Bad GatewayIt usually targets proxy and upstream connections; 504 is more compatible with long wait and queue.

Step 6: Database and related services

The program may be running but will be waiting for the external database, Redis, storage, DNS, or API. Check connection, latency, pool, and authentication by log. Change the password or certificate dependency after deploying the RIGI scenario. Do not print credential for testing in the command line or log.

The full disc and the inode are complete.

The disk can stop the database write, session, log, and startup by 100%. The inode may also be finished with the external space. Before deleting, identify the consumption and consider the file opened by the process.

RAM, Swap and OOM Killer.

If the kernel has killed the critical process memory release, restart only repeats the cycle. See kernel events, the final pre-consumption, and the memory of each process.

CPU, load and I/O.

CPU 100%, high load and I/O wait are three different situations. Find process user, start time and workload. Backup, compression, crawler, query or timed job may be overlapping with traffic. Killing an unknown process can impair write; first define its role and effect.

Is the recent deployment a definite factor?

Compare the image version, commit, migration, config, and deploy time to the start of the incident. Correlation is not the definite cause, but makes rollback a verifiable option. Rollback is only safe when the new schema and data are compatible with the previous version. Manual deployment without an accurate manifest does not return reliable.

Restore service with minimal change.

  1. Record the evidence and current status.
  2. Identify the faulty layer and user effect.
  3. If the recent change is clear and rollback is safe, make a controlled return.
  4. Otherwise, only run the faulty service after you understand the failure start/reload.
  5. Test the user's technical health and main journey.
  6. Consider the metric and log to return the error below.

Uploading is not the end of the process incident. DNS, TLS, dynamic page, login and critical operations must be externally verified. If data is left in queue or callback is repeated, reconciliation is required.

What things are making the situation worse?

  • Restart simultaneously the web, the app and the database
  • Clear logs and cache without a hypothesis.
  • Turn off the firewall or security for public testing.
  • Delete large files without identifying the owner or backup.
  • Increase all timeouts and limits.
  • Change DNS and server at the same time.
  • Load test on the half-damaged system.
  • Statement of cause before evidence is confirmed

After the resuscitation: Root Cause Analysis

Record timeline, trigger, technical cause, aggravating factors, reason for delay in detection and reason for delay in recovery. RCA is not to blame for finding; it must create a proprietary action with a deadline.

Practical Prevention

  • External monitoring of DNS, TLS and HTTP
  • Metric sources, queue, dependency and error rate
  • Alert with Responsibility, Severity and Runbook.
  • Backup recovered and RPO/RTO tested
  • Deploy a step with rollback
  • Capacity and load test controlled in staging
  • Emergency console and separate accesses
  • Registering changes and maintaining a copy of the config without secret

Checklist of early configuration of LinuxBaseline covers access, backup and monitoring to prevent incidents.

When do you need special assistance?

If it is completely repeated, the cause disappears after restarting or multiple layers of infrastructure are involved, manipulation of the production is more likely to result in evidence loss.Monthly management of the serverIt can respond to incident, correlation logs, RCA and preventive action in a traceable process.

Common Questions

What's the first thing to do when the site goes down?

Record the time and effect range and measure the DNS/TLS/HTTP from outside; save log snapshots and resources before restarting.

If the server ping, is it safe?

No. Ping represents only a portion of the network access and does not verify the integrity of the port, TLS, web server, application, and database.

Is rebooting the server the right solution?

Only when the cause or runbook justifies it. Reboot blinds the evidence and may also reveal unstarted services.

Why does the site go down again after restarting?

The root cause, such as memory leakage, full disk, heavy job or spoiled dependency, remains; compare metrics and logs before any particulars.

How to Fix a 504 Gateway Timeout Without Raising Timeouts Blindly
Fix 504 by identifying the payment gatewayway builder error, upstream time, PHP row, query database, external API and limit each layer; not increasing arbitrary timeout.