Skip to content

How to Fix a 504 Gateway Timeout Without Raising Timeouts Blindly

Fix 504 by identifying the payment gatewayway builder error, upstream time, PHP row, query database, external API and limit each layer; not increasing arbitrary timeout.

Author Bipida Editorial Team Published
Share this article

It's a mistake.504 Gateway TimeoutThis means that a gateway has not received an upstream response in full at the time it is set. The payment gatewayway can be a CDN, load balancer, Nginx or proxy within the platform. The multiple timeout increase may indicate the error page later, but the Query is locked, the API is outgoing, or the PHP-FPM full row is not speeding up, and even occupies more capacity.

Quick answer:First, determine which layer generates 504 and after a few seconds. Follow the request slow with a timestamp or request ID in the access/error log, separate the proxy and upstream time, and then check the PHP, database, job, or API associated with it.

Signs that cut the pathway for diagnosis

  • All pages after a period of time become almost constant 504.
  • Only the report, import, search or checkout is wrong.
  • The error is only seen at peak times.
  • You get the answer from the origin, but not from the CDN, or vice versa.
  • The request for 504 continues to be backend.
  • Deploy or change the database before the error starts

Fixed time often refers to the timeout of a layer. A particular path error is closer to the logic of the same endpoint or dependency. A peak error can be queue, connection pool, or saturation. These are hypotheses; they must be confirmed by logs and metrics.

Which Gateway built the 504 page?

Check the error page, response header, Ray/Request ID, and logs. Nginx may still be waiting but the CDN will timeout earlier. Or the load balancer is healthy and Nginx is not responding from the internal app.

Create a real timeline.

  1. Record the time of the request start with the timezone.
  2. Measure the time until the 504 display.
  3. Follow the request ID from edge to program.
  4. Separate the connection time and the upstream response time.
  5. Check the end or continuation of the backend processing after the client is interrupted.

If the request works in the program layer for 90 seconds but the proxy stops at 60 seconds, the next question is not what number to do 120; we first need to figure out whether the interactive operation really should take 90 seconds or whether it's better to turn it into a background job.

What does the Nginx log say?

Message from the police.upstream timed outAccess logs better represent the slow layer if the request has time and upstream response time/status. The empty or multiple upstream values in the log must also be interpreted correctly.

Check the status of resources without creating load

uptime
free -h
df -h
df -i
ss -s
systemctl --failed

This snapshot is just the status of the moment and does not replace the historical metric. High load is not always high CPU; waiting task I/O also increases load. RAM is not necessarily a crisis in Linux; see available, swap, OOM and process behavior together.

PHP-FPM and the Satisfied Workers

If all workers are busy with slow requests, the new request will stay in line and will be timeout before processing begins. Do not simply increase the number of workers. Measure the average and maximum memory of each process, allowable RAM, CPU, and request length. Slow log control can show a long-track stack; manage sensitive data.

Query or database lock

Heavy reporting, inappropriate indexing, lock or connection pool filling can stop the request. Check the slow query log and database status at the time of the occurrence. Activating large-scale logging on production requires limited disk capacity and time. Do not run a malicious query or delete operation for test and have backup and rollback before changing the schema.

For the WordPress,The Slow Query search guide.It helps separate PHP time from SQL time. The profiler plugin should also be used short-term and overhead metering.

API and external service

Checkout, texting, taxation, transportation, or authentication may be waiting for another API. Connection, reading, and retry timeout should be limited and visible. Retry without idempotency can run a financial operation twice.

DNS and network.

Resolve, packet loss, network path or connection establishment can also be time consuming. Testing should be done from the same host/container and network; the resulting administrator laptop is not necessarily representative of the production path. Consider IPv4 and IPv6, proxy environment and DNS resolver activated separately.

Separate long work from interactive HTTP

Large export, file generation, sync, or image processing is better assigned to a traceable status background job. It provides a quick response to a job ID and the separate worker performs operations with controlled retry. This architecture change is more stable than long timeout and maintains a connection for a few minutes.

Multi-layered timeout

Browser/client, CDN, load balancer, Nginx, application server, PHP, and database each have different timeout. The ceilings must be aligned based on SLA endpoint and failure budget. The internal ceiling should usually allow for controlled error to be recorded before the external layer is cut, but the exact number cannot be prescribed without architecture and workload.

What happens to the backend after the payment gatewayway is shut down?

504 display does not necessarily mean a stop in the backend. Depending on the runtime and how the client is detected, the query, job, or call to the payment gateway may continue. Therefore, the user should not repeat the financial operation or order record without checking the result. Status-changing endpoints must have a unique identifier and idempotent behavior to prevent controlled retry from recording repeat.

In the incident, check whether the timeout requests still hold the process or connection. If so, the entry of new retries will increase the backlog. Temporary traffic restriction or disabling of a feature should be done by evaluating the business impact; shutting down the entire system or mass kill of processes can result in a semi-complete and inconsistent data transaction.

Connect time to response time

Delay in connecting usually takes a look at the DNS, network, listener, and pool of connections; delayed response after connections is more related to queue and program execution or dependency. Record these two in separate metrics and logs. A total number time alone does not tell the time before the program is consumed or within the logic of the request.

When is the timeout increase appropriate?

When the operation is healthy and limited, it is deliberately long; there are sufficient resources; progress or client interruption behavior is obvious; and the current ceiling is lower than the business's acceptable time. Even then, just change the directive for the actual timeout stage and then track the error rate, concurrency, and memory.

The order of the problem.

  1. Enter the request sample, time and ID.
  2. Find the 504 manufacturing layer and the timing ceiling.
  3. Separate upstream time and queue from log/metric.
  4. Check PHP, databases, APIs, DNS and sources based on evidence.
  5. Remove the cause of prolonged or saturated processing.
  6. If long time is inherent, design an async or timeout document architecture.
  7. Test in staging and then with limited rollout.
  8. Control the main business direction and error rate after the change.

Common errors when filing 504

  • Increases all timeout without knowingly layering.
  • Restart before keeping log and metric
  • Unlimited load test on the production
  • Increase worker without memory budget
  • Ignoring the external database or API lock
  • Retry is an idempotent transaction.
  • One knew 502 and 504.
  • Judging only by the average latency.

Preventive monitoring is needed.

Monitor p50/p95/p99, rate 504, request and upstream time, queue, active worker, connection database, dependency latency and saturation of resources. Alert must be activated before the timeout is taken.

When do you need special assistance?

If 504 is repeated in Checkout, panel, or important API and the time source is ambiguous across multiple layers, the increased timeout increases the risk of backlog.Monthly management of the serverYou can track the time when Nginx, PHP-FPM, databases and dependencies were put together and targeted the collar.

Common Questions

Is 504 a weak server?

Not always. Locked queries, external APIs, DNS, config timeout, or rows can also be the cause; sources are just one of the assumptions.

What's the difference between 504 and 502?

504 refers more to the expiration of the response time; 502 is usually an upstream failure or unreliable response.

Why do we only have 504 during rush hour?

The probability of saturation worker, connection or dependency is higher. Match the capacity metric and latency to the request rate.

Did they get up?proxy_read_timeoutIs that enough?

Only if the same phase is timeout and the duration of the operation is accepted and its capacity is calculated; otherwise it will display the signal later.

How to Fix a 502 Bad Gateway Error in Nginx
To fix 502 Nginx, match the error time with the error log and check the upstream status, PHP-FPM, socket, permission, worker capacity and proxy chain.