Debug guide · you're seeing
502 Bad Gatewayconnect() failed (111: Connection refused) while connecting to upstreamupstream prematurely closed connection while reading response header from upstreamconnect() to unix:/run/php/php8.3-fpm.sock failed (13: Permission denied) while connecting to upstreamupstream sent too big header while reading response header from upstreamrecv() failed (104: Connection reset by peer) while reading response header from upstream
nginx 502 Bad Gateway: Causes and Fixes by Error Log
Fix nginx 502 Bad Gateway by matching the error log: connection refused, prematurely closed connection, php-fpm socket permissions, too big header, keepalive.
On this page
- What it means
- Who sent it?
- Fix it, matched to the log line
- connect() failed (111: Connection refused) while connecting to upstream
- connect() to unix:/run/php/php8.3-fpm.sock failed (13: Permission denied) or (2: No such file or directory)
- upstream prematurely closed connection while reading response header from upstream
- upstream sent too big header while reading response header from upstream
- Intermittent 502: keepalive mismatch
- no live upstreams while connecting to upstream
- Reproduce and verify
- Related
TL;DR: nginx could not get a valid response from the upstream. Open the nginx error log and match the line after
upstream:Connection refusedmeans nothing is listening,Permission deniedis a socket or SELinux issue,prematurely closed connectionmeans the app crashed or timed out internally,too big headerisproxy_buffer_size.
What it means
502 is a gateway reporting a bad answer from further along the chain (RFC 9110 section 15.6.3). In nginx it is generated when the connection to the upstream fails or the upstream response is invalid:
HTTP/1.1 502 Bad Gateway
Server: nginx/1.27.0
Content-Type: text/html
The access log shows the 502 but not the reason. The reason is in the error log, one line per failure, and everything useful follows the word upstream:
2026/10/04 11:20:05 [error] 31#31: *912 connect() failed (111: Connection refused) while connecting to upstream, client: 203.0.113.10, server: example.com, request: "GET / HTTP/1.1", upstream: "http://127.0.0.1:3000/", host: "example.com"
Start there. Everything below is organized by that log line.
Who sent it?
Server: nginxin the response and a line in your nginx error log with the same timestamp: this page applies.Server: awselb/2.0(ALB): the load balancer, not nginx. Look at its access logelb_status_codeandtarget_status_code. A 502 withtarget_status_codeof-means the target closed the connection or sent a malformed response, often a keepalive race (below).- Cloudflare with an error page titled “Bad gateway” and a
CF-Ray: Cloudflare got an invalid response from the origin. Errors 520 to 523 are Cloudflare’s finer-grained cousins (unknown error, web server is down, connection timed out, origin unreachable). - No nginx log line at all: the request did not reach this nginx. Something earlier generated the 502.
Fix it, matched to the log line
connect() failed (111: Connection refused) while connecting to upstream
Nothing is accepting connections at the upstream: address. The process crashed, never started, listens on a different port or interface, or nginx is pointing at the wrong place.
# Is anything listening on the port nginx is using?
ss -ltnp | grep ':3000'
# Talk to the upstream directly, bypassing nginx
curl -i http://127.0.0.1:3000/
# Why did the app stop? (systemd, then the kernel's OOM killer)
journalctl -u myapp --since '10 min ago'
dmesg | grep -i 'killed process'
Common causes: the app is bound to localhost inside a container while nginx connects to the container IP (bind to 0.0.0.0), the app crashed on boot after a deploy, or the process was OOM-killed.
On RHEL, CentOS and Fedora with SELinux enforcing, nginx is not allowed to open network connections by default and the log shows (13: Permission denied) instead. Allow it with setsebool -P httpd_can_network_connect 1.
DNS is a separate trap. proxy_pass http://backend.internal; resolves the name once when nginx starts or reloads. If the IP changes (Docker, Kubernetes, ECS, load balancers with rotating addresses), nginx keeps connecting to the old address and gets refused or timed out. Make nginx re-resolve:
resolver 127.0.0.11 valid=10s; # your DNS server; 127.0.0.11 is Docker's embedded DNS
set $upstream http://backend.internal:3000;
proxy_pass $upstream;
Using a variable in proxy_pass forces runtime resolution through resolver.
connect() to unix:/run/php/php8.3-fpm.sock failed (13: Permission denied) or (2: No such file or directory)
The nginx worker user cannot open the PHP-FPM socket, or the socket path in fastcgi_pass does not match the pool’s listen setting.
; /etc/php/8.3/fpm/pool.d/www.conf
listen = /run/php/php8.3-fpm.sock
listen.owner = www-data
listen.group = www-data
listen.mode = 0660
# nginx.conf: the user directive must match listen.owner or listen.group
user www-data;
location ~ \.php$ {
include fastcgi_params;
fastcgi_pass unix:/run/php/php8.3-fpm.sock;
}
After a PHP version upgrade the socket file name often changes (php8.2-fpm.sock to php8.3-fpm.sock) and fastcgi_pass still points at the old one, which gives error 2. Check with ls -l /run/php/.
upstream prematurely closed connection while reading response header from upstream
nginx connected and sent the request, and the upstream closed the connection without sending a complete response. The upstream process died or was killed mid-request:
- A worker crash (segfault, unhandled exception, out-of-memory kill).
- A worker killed by its own supervisor for taking too long. Gunicorn kills a worker after
--timeout(30 seconds by default) and nginx logs this line, not a timeout. PHP-FPM’srequest_terminate_timeoutand PHP’smax_execution_timedo the same. - A deploy or restart that stops workers while requests are in flight.
Look at the upstream’s logs at the same timestamp; the nginx line only tells you the connection was cut. For slow-endpoint cases, see nginx 504 and keep nginx’s timeouts higher than the application’s, so the application can answer with a proper error before nginx gives up.
recv() failed (104: Connection reset by peer) while reading response header from upstream is the same family: the upstream sent a TCP RST. Check the keepalive section next if it is intermittent, and see ERR_CONNECTION_RESET for finding which hop sent the reset.
upstream sent too big header while reading response header from upstream
The status line and headers of the upstream response must fit in a single buffer, sized by proxy_buffer_size, which defaults to one memory page (4k or 8k depending on the platform; ingress-nginx sets 4k). Large cookies, many Set-Cookie headers, long Link preload headers or huge redirect URLs overflow it, and the client gets a 502 only on those routes (login is the usual one).
location / {
proxy_pass http://app;
proxy_buffer_size 16k; # first part of the response: headers
proxy_buffers 8 16k;
proxy_busy_buffers_size 32k; # must be >= proxy_buffer_size and < total buffers minus one buffer
}
For FastCGI upstreams the equivalents are fastcgi_buffer_size, fastcgi_buffers and fastcgi_busy_buffers_size. On ingress-nginx use the nginx.ingress.kubernetes.io/proxy-buffer-size: "16k" annotation. If 16k is not enough, the real fix is smaller cookies and headers.
Intermittent 502: keepalive mismatch
When a proxy reuses an idle connection at the exact moment the other side closes it, the request is written onto a dead connection and the proxy sees a reset. It looks random and clusters under steady, low traffic.
- nginx to upstream: if you enable upstream keepalive, the application’s idle timeout must be longer than nginx’s. nginx closes its own idle upstream connections after
keepalive_timeout(60 seconds by default in the upstream block). Node’shttp.Servercloses idle connections after 5 seconds by default, so nginx regularly reuses connections Node has just closed. - ALB to target: the load balancer’s idle timeout is 60 seconds by default. A backend that closes idle connections sooner produces 502s from the ALB. Set the application keepalive higher than the ALB’s idle timeout.
// Node: keep idle connections open longer than the ALB (60s) or nginx upstream keepalive
const server = app.listen(3000)
server.keepAliveTimeout = 65_000 // ms; Node's default is 5000, so set it explicitly
server.headersTimeout = 66_000 // must be greater than keepAliveTimeout
upstream app {
server 127.0.0.1:3000;
keepalive 32;
keepalive_timeout 30s; # keep this below the application's idle timeout
}
server {
location / {
proxy_pass http://app;
proxy_http_version 1.1; # required for upstream keepalive
proxy_set_header Connection ""; # do not forward "close"
}
}
no live upstreams while connecting to upstream
Every server in the upstream block is marked failed after max_fails errors (default 1) within fail_timeout (default 10 seconds), and nginx stops trying them for that window. Fix the upstream, or raise max_fails if blips are expected. A single upstream server is never marked down, so this message means you have more than one.
Reproduce and verify
# 1. Does the app answer directly?
curl -si http://127.0.0.1:3000/ | head -n 5
# 2. Does it answer through nginx, with the same Host?
curl -si http://127.0.0.1/ -H 'Host: example.com' | head -n 5
# 3. Watch the error log while you retry
tail -f /var/log/nginx/error.log
# 4. Validate and reload after config edits
nginx -t && nginx -s reload
For a header-size problem, measure the response headers the upstream sends: curl -sD - -o /dev/null http://127.0.0.1:3000/login | wc -c. If it is near or above proxy_buffer_size, that is your answer.
Related
- 502 Bad Gateway explains the status and compares it to 500 and 504.
- 504 Gateway Timeout: the upstream was reachable but too slow.
- upstream sent too big header: the 502 caused by response headers that overflow
proxy_buffer_size. - 503 Service Unavailable is the status an upstream should return when it is overloaded or draining, instead of dropping the connection.
- Connection and Keep-Alive cover the persistent connection behavior behind the intermittent cases.
Frequently asked questions
What does a 502 Bad Gateway mean in nginx?
nginx acted as a proxy, tried to get a response from the upstream, and either could not connect or received something it could not use: a closed connection, a reset, or a malformed response. It is the proxy reporting a failure behind it, so the cause is in the upstream or the path between nginx and the upstream, not in the browser.
Where do I find out why nginx returned 502?
The error log, usually /var/log/nginx/error.log, or the ingress controller pod logs in Kubernetes. The access log only shows the 502 status; the error log has the specific reason after the word upstream, such as Connection refused, prematurely closed connection, or too big header.
Why do I get 502 only on some requests, or right after a deploy?
Intermittent 502 usually means the upstream restarts, crashes or closes keepalive connections while nginx still reuses them. During a rolling deploy, a worker that is shut down mid-request produces prematurely closed connection. A race between an application keepalive timeout and the proxy reusing the idle connection produces occasional connection reset errors.
What is the difference between 502 and 504?
A 502 means the upstream was unreachable or answered invalidly, usually quickly. A 504 means the upstream accepted the request but did not answer in time. If a worker is killed because it ran too long, the symptom is 502, not 504, because the connection is cut rather than timed out.
Why is my 502 only on large responses or logins?
The upstream response headers are bigger than nginx's proxy_buffer_size, which defaults to one memory page (4k or 8k). Large Set-Cookie headers, long redirects with big query strings, or verbose debug headers trigger upstream sent too big header. Raise proxy_buffer_size or shrink the headers.
Sources
Related
nginx 504 Gateway Timeout: Fix Upstream Timed Out (110)
Fix nginx 504 Gateway Time-out and 'upstream timed out (110)': proxy_read_timeout, fastcgi_read_timeout, ALB idle timeout, and Cloudflare 524 compared.
HTTP 502 Bad Gateway: nginx, ALB and Cloudflare Fixes
Fix 502 Bad Gateway: decode nginx error-log lines, php-fpm sockets, ALB keep-alive mismatches and Cloudflare 502 vs 52x, with curl checks.
HTTP 503 Service Unavailable: Causes, Fixes and Retry-After
Fix HTTP 503 Service Unavailable: nginx no live upstreams, Kubernetes endpoints, ALB healthy hosts, Cloudflare, and a maintenance page with Retry-After.
504 Gateway Timeout: nginx, ALB and Cloudflare Fixes
Fix 504 Gateway Timeout: nginx proxy_read_timeout (60s default), ALB 60s idle, API Gateway 29s, Cloudflare 524 at 125s, with error-log strings and curl timing.