# X-Robots-Tag Header: noindex for PDFs and Staging

> X-Robots-Tag sends noindex and snippet rules as an HTTP header for PDFs, images and staging sites. nginx, Apache and Cloudflare config, and the robots.txt trap.

Source: https://howhttpworks.com/headers/x-robots-tag
Last reviewed: 2026-10-04

> **TL;DR:** `X-Robots-Tag: noindex` tells search engines not to list a URL, and it works on PDFs, images and other files where a meta tag is impossible. The crawler has to be allowed to fetch the URL to see it, so never combine it with a `robots.txt` Disallow on the same path.

## Syntax

```http
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex, nofollow
```

The value is a comma-separated list of rules. You can send several headers, and you can scope a rule to one crawler by prefixing its user agent token:

```http
X-Robots-Tag: googlebot: noindex
X-Robots-Tag: otherbot: noindex, nofollow
```

Rules without a user agent apply to every crawler that honours the header. Google documents that the header name, user agent names and values are not case sensitive. When rules conflict, the more restrictive one wins.

## Directives Google supports

Per Google Search Central's robots meta tag specification:

| Rule | Effect |
| --- | --- |
| `all` | No restrictions. This is the default. |
| `noindex` | Do not show the page, media or resource in search results. |
| `nofollow` | Do not follow links on the page. |
| `none` | Shorthand for `noindex, nofollow`. |
| `nosnippet` | Show no text snippet or video preview. |
| `indexifembedded` | Allow indexing of the content when it is embedded in another page through an iframe or similar, even though the page carries `noindex`. It only takes effect together with `noindex`. |
| `max-snippet: [number]` | Limit the text snippet to that many characters. `0` means none, `-1` means no limit. |
| `max-image-preview: [setting]` | `none`, `standard` or `large`. |
| `max-video-preview: [number]` | Limit video previews to that many seconds. `0` allows a static image only, `-1` is unlimited. |
| `notranslate` | Do not offer translation of the page in results. |
| `noimageindex` | Do not index images on the page. |
| `unavailable_after: [date]` | Stop showing the page after the date. RFC 822, RFC 850 and ISO 8601 formats are accepted. |

Other search engines document their own sets. Check each vendor before relying on a directive outside this list.

## When to use it

- **Files with no HTML head.** PDFs, Word and Excel downloads, images, and video. A meta tag cannot exist in them.
- **Staging, preview and internal hosts.** Send `noindex, nofollow` for every response from the host. It is quicker to apply than editing templates.
- **Pagination, filter and search-result URLs** that you want crawled for links but not indexed: `noindex, follow`.
- **Expiring content.** `unavailable_after: 31 Dec 2027 23:59:59 GMT` for a promotion page.

If the content is confidential, use authentication. `noindex` is a request to crawlers, not access control.

## Configuration

### nginx

```nginx
# Every PDF and Word document on the site
location ~* \.(pdf|docx?|xlsx?)$ {
    add_header X-Robots-Tag "noindex, nofollow" always;
    try_files $uri =404;
}

# Whole staging host
server {
    server_name staging.example.com;
    add_header X-Robots-Tag "noindex, nofollow" always;
}
```

Two nginx behaviours cause most "it works on one URL and not another" bugs. Without `always`, `add_header` only applies to 200, 201, 204, 206 and the 3xx codes, so a 404 or 500 page goes out without it. And `add_header` inherits from the enclosing level only if the current level has no `add_header` of its own: put a `Cache-Control` `add_header` in a `location` and the server-level `X-Robots-Tag` disappears for that location. nginx 1.29.3 added `add_header_inherit merge` to change that. On older versions, repeat the header in each block.

### Apache

```apache
<FilesMatch "\.(pdf|docx?|xlsx?)$">
  Header set X-Robots-Tag "noindex, nofollow"
</FilesMatch>
```

This is Google's own example pattern for PDFs and requires `mod_headers`. Use `Header always set` if error responses need it too.

### Cloudflare

For a site proxied through Cloudflare, add a Response Header Transform Rule (Rules, Transform Rules, Modify Response Header): set a static header named `X-Robots-Tag` with value `noindex, nofollow`, matched by a custom filter expression such as:

```text
ends_with(http.request.uri.path, ".pdf") or http.host eq "staging.example.com"
```

For Cloudflare Pages static assets, use a `_headers` file in the build output. Cloudflare documents `X-Robots-Tag: noindex` for preview deployments as a use case:

```text
https://:project.pages.dev/*
  X-Robots-Tag: noindex
```

`_headers` rules do not apply to responses generated by Pages Functions. Set the header in the Function code for those.

### Express

```javascript
app.use('/downloads', (req, res, next) => {
  res.set('X-Robots-Tag', 'noindex, nofollow')
  next()
})
```

## The robots.txt trap

`robots.txt` controls crawling. `X-Robots-Tag` is read from the response of a crawl. If a URL is disallowed, the crawler never makes the request, so it never sees the header. Google's documentation says it plainly: if the page is blocked by `robots.txt` the crawler never sees the `noindex` rule and the page can still appear in search results. Google also does not support a `noindex` line inside `robots.txt`.

So the sequence for removing something is:

1. Remove the Disallow rule for that path.
2. Serve `X-Robots-Tag: noindex` on the URL.
3. Request a recrawl in Search Console's URL Inspection tool, or wait.
4. Only after it has dropped out, if you want to stop crawling, add the Disallow back.

## Verify

```bash
curl -sI https://example.com/files/report.pdf | grep -i x-robots-tag
```

Run it against the canonical, final URL. Then run it against a 404 path on the same host to see whether the header survives error responses. Search Console's URL Inspection shows whether Google saw `noindex` on its last fetch.

## Related

- [Content-Type](https://howhttpworks.com/headers/content-type), the header that tells crawlers a URL is a PDF or an image in the first place
- [Cache-Control](https://howhttpworks.com/headers/cache-control): crawlers recrawl on their own schedule, but a long `max-age` at your CDN can keep serving an old header
