# nginx 502 Bad Gateway: Causes and Fixes by Error Log

> Fix nginx 502 Bad Gateway by matching the error log: connection refused, prematurely closed connection, php-fpm socket permissions, too big header, keepalive.

Source: https://howhttpworks.com/debug/nginx-502-bad-gateway
Last reviewed: 2026-10-04

Error messages this page covers:
- `502 Bad Gateway`
- `connect() failed (111: Connection refused) while connecting to upstream`
- `upstream prematurely closed connection while reading response header from upstream`
- `connect() to unix:/run/php/php8.3-fpm.sock failed (13: Permission denied) while connecting to upstream`
- `upstream sent too big header while reading response header from upstream`
- `recv() failed (104: Connection reset by peer) while reading response header from upstream`

> **TL;DR:** nginx could not get a valid response from the upstream. Open the nginx error log and match the line after `upstream`: `Connection refused` means nothing is listening, `Permission denied` is a socket or SELinux issue, `prematurely closed connection` means the app crashed or timed out internally, `too big header` is `proxy_buffer_size`.

## What it means

502 is a gateway reporting a bad answer from further along the chain (RFC 9110 section 15.6.3). In nginx it is generated when the connection to the upstream fails or the upstream response is invalid:

```http
HTTP/1.1 502 Bad Gateway
Server: nginx/1.27.0
Content-Type: text/html
```

The access log shows the 502 but not the reason. The reason is in the error log, one line per failure, and everything useful follows the word `upstream`:

```text
2026/10/04 11:20:05 [error] 31#31: *912 connect() failed (111: Connection refused) while connecting to upstream, client: 203.0.113.10, server: example.com, request: "GET / HTTP/1.1", upstream: "http://127.0.0.1:3000/", host: "example.com"
```

Start there. Everything below is organized by that log line.

## Who sent it?

- `Server: nginx` in the response and a line in your nginx error log with the same timestamp: this page applies.
- `Server: awselb/2.0` (ALB): the load balancer, not nginx. Look at its access log `elb_status_code` and `target_status_code`. A 502 with `target_status_code` of `-` means the target closed the connection or sent a malformed response, often a keepalive race (below).
- Cloudflare with an error page titled "Bad gateway" and a `CF-Ray`: Cloudflare got an invalid response from the origin. Errors 520 to 523 are Cloudflare's finer-grained cousins (unknown error, web server is down, connection timed out, origin unreachable).
- No nginx log line at all: the request did not reach this nginx. Something earlier generated the 502.

## Fix it, matched to the log line

### `connect() failed (111: Connection refused) while connecting to upstream`

Nothing is accepting connections at the `upstream:` address. The process crashed, never started, listens on a different port or interface, or nginx is pointing at the wrong place.

```bash
# Is anything listening on the port nginx is using?
ss -ltnp | grep ':3000'

# Talk to the upstream directly, bypassing nginx
curl -i http://127.0.0.1:3000/

# Why did the app stop? (systemd, then the kernel's OOM killer)
journalctl -u myapp --since '10 min ago'
dmesg | grep -i 'killed process'
```

Common causes: the app is bound to `localhost` inside a container while nginx connects to the container IP (bind to `0.0.0.0`), the app crashed on boot after a deploy, or the process was OOM-killed.

On RHEL, CentOS and Fedora with SELinux enforcing, nginx is not allowed to open network connections by default and the log shows `(13: Permission denied)` instead. Allow it with `setsebool -P httpd_can_network_connect 1`.

DNS is a separate trap. `proxy_pass http://backend.internal;` resolves the name once when nginx starts or reloads. If the IP changes (Docker, Kubernetes, ECS, load balancers with rotating addresses), nginx keeps connecting to the old address and gets refused or timed out. Make nginx re-resolve:

```nginx
resolver 127.0.0.11 valid=10s;      # your DNS server; 127.0.0.11 is Docker's embedded DNS
set $upstream http://backend.internal:3000;
proxy_pass $upstream;
```

Using a variable in `proxy_pass` forces runtime resolution through `resolver`.

### `connect() to unix:/run/php/php8.3-fpm.sock failed (13: Permission denied)` or `(2: No such file or directory)`

The nginx worker user cannot open the PHP-FPM socket, or the socket path in `fastcgi_pass` does not match the pool's `listen` setting.

```ini
; /etc/php/8.3/fpm/pool.d/www.conf
listen = /run/php/php8.3-fpm.sock
listen.owner = www-data
listen.group = www-data
listen.mode = 0660
```

```nginx
# nginx.conf: the user directive must match listen.owner or listen.group
user www-data;
```

```nginx
location ~ \.php$ {
    include fastcgi_params;
    fastcgi_pass unix:/run/php/php8.3-fpm.sock;
}
```

After a PHP version upgrade the socket file name often changes (`php8.2-fpm.sock` to `php8.3-fpm.sock`) and `fastcgi_pass` still points at the old one, which gives error 2. Check with `ls -l /run/php/`.

### `upstream prematurely closed connection while reading response header from upstream`

nginx connected and sent the request, and the upstream closed the connection without sending a complete response. The upstream process died or was killed mid-request:

- A worker crash (segfault, unhandled exception, out-of-memory kill).
- A worker killed by its own supervisor for taking too long. Gunicorn kills a worker after `--timeout` (30 seconds by default) and nginx logs this line, not a timeout. PHP-FPM's `request_terminate_timeout` and PHP's `max_execution_time` do the same.
- A deploy or restart that stops workers while requests are in flight.

Look at the upstream's logs at the same timestamp; the nginx line only tells you the connection was cut. For slow-endpoint cases, see [nginx 504](https://howhttpworks.com/debug/nginx-504-gateway-timeout) and keep nginx's timeouts higher than the application's, so the application can answer with a proper error before nginx gives up.

`recv() failed (104: Connection reset by peer) while reading response header from upstream` is the same family: the upstream sent a TCP RST. Check the keepalive section next if it is intermittent, and see [ERR_CONNECTION_RESET](https://howhttpworks.com/debug/err-connection-reset) for finding which hop sent the reset.

### `upstream sent too big header while reading response header from upstream`

The status line and headers of the upstream response must fit in a single buffer, sized by `proxy_buffer_size`, which defaults to one memory page (4k or 8k depending on the platform; ingress-nginx sets `4k`). Large cookies, many `Set-Cookie` headers, long `Link` preload headers or huge redirect URLs overflow it, and the client gets a 502 only on those routes (login is the usual one).

```nginx
location / {
    proxy_pass http://app;
    proxy_buffer_size        16k;   # first part of the response: headers
    proxy_buffers            8 16k;
    proxy_busy_buffers_size  32k;   # must be >= proxy_buffer_size and < total buffers minus one buffer
}
```

For FastCGI upstreams the equivalents are `fastcgi_buffer_size`, `fastcgi_buffers` and `fastcgi_busy_buffers_size`. On ingress-nginx use the `nginx.ingress.kubernetes.io/proxy-buffer-size: "16k"` annotation. If 16k is not enough, the real fix is smaller cookies and headers.

### Intermittent 502: keepalive mismatch

When a proxy reuses an idle connection at the exact moment the other side closes it, the request is written onto a dead connection and the proxy sees a reset. It looks random and clusters under steady, low traffic.

- nginx to upstream: if you enable upstream keepalive, the application's idle timeout must be longer than nginx's. nginx closes its own idle upstream connections after `keepalive_timeout` (60 seconds by default in the upstream block). Node's `http.Server` closes idle connections after 5 seconds by default, so nginx regularly reuses connections Node has just closed.
- ALB to target: the load balancer's idle timeout is 60 seconds by default. A backend that closes idle connections sooner produces 502s from the ALB. Set the application keepalive higher than the ALB's idle timeout.

```javascript
// Node: keep idle connections open longer than the ALB (60s) or nginx upstream keepalive
const server = app.listen(3000)
server.keepAliveTimeout = 65_000   // ms; Node's default is 5000, so set it explicitly
server.headersTimeout = 66_000     // must be greater than keepAliveTimeout
```

```nginx
upstream app {
    server 127.0.0.1:3000;
    keepalive 32;
    keepalive_timeout 30s;   # keep this below the application's idle timeout
}

server {
    location / {
        proxy_pass http://app;
        proxy_http_version 1.1;           # required for upstream keepalive
        proxy_set_header Connection "";   # do not forward "close"
    }
}
```

### `no live upstreams while connecting to upstream`

Every server in the `upstream` block is marked failed after `max_fails` errors (default 1) within `fail_timeout` (default 10 seconds), and nginx stops trying them for that window. Fix the upstream, or raise `max_fails` if blips are expected. A single upstream server is never marked down, so this message means you have more than one.

## Reproduce and verify

```bash
# 1. Does the app answer directly?
curl -si http://127.0.0.1:3000/ | head -n 5

# 2. Does it answer through nginx, with the same Host?
curl -si http://127.0.0.1/ -H 'Host: example.com' | head -n 5

# 3. Watch the error log while you retry
tail -f /var/log/nginx/error.log

# 4. Validate and reload after config edits
nginx -t && nginx -s reload
```

For a header-size problem, measure the response headers the upstream sends: `curl -sD - -o /dev/null http://127.0.0.1:3000/login | wc -c`. If it is near or above `proxy_buffer_size`, that is your answer.

## Related

- [502 Bad Gateway](https://howhttpworks.com/status-codes/502) explains the status and compares it to 500 and 504.
- [504 Gateway Timeout](https://howhttpworks.com/debug/nginx-504-gateway-timeout): the upstream was reachable but too slow.
- [upstream sent too big header](https://howhttpworks.com/debug/nginx-upstream-sent-too-big-header): the 502 caused by response headers that overflow `proxy_buffer_size`.
- [503 Service Unavailable](https://howhttpworks.com/status-codes/503) is the status an upstream should return when it is overloaded or draining, instead of dropping the connection.
- [Connection](https://howhttpworks.com/headers/connection) and [Keep-Alive](https://howhttpworks.com/headers/keep-alive) cover the persistent connection behavior behind the intermittent cases.
