Response header · non-standard
< X-Robots-Tag: noindex, nofollowX-Robots-Tag Header: noindex for PDFs and Staging
X-Robots-Tag sends noindex and snippet rules as an HTTP header for PDFs, images and staging sites. nginx, Apache and Cloudflare config, and the robots.txt trap.
- Direction
- Response
- Category
- SEO
- Spec
- MDN reference
- Status
- Non-standard: Crawler directive; see Google Search Central
On this page
TL;DR:
X-Robots-Tag: noindextells search engines not to list a URL, and it works on PDFs, images and other files where a meta tag is impossible. The crawler has to be allowed to fetch the URL to see it, so never combine it with arobots.txtDisallow on the same path.
Syntax
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex, nofollow
The value is a comma-separated list of rules. You can send several headers, and you can scope a rule to one crawler by prefixing its user agent token:
X-Robots-Tag: googlebot: noindex
X-Robots-Tag: otherbot: noindex, nofollow
Rules without a user agent apply to every crawler that honours the header. Google documents that the header name, user agent names and values are not case sensitive. When rules conflict, the more restrictive one wins.
Directives Google supports
Per Google Search Central’s robots meta tag specification:
| Rule | Effect |
|---|---|
all | No restrictions. This is the default. |
noindex | Do not show the page, media or resource in search results. |
nofollow | Do not follow links on the page. |
none | Shorthand for noindex, nofollow. |
nosnippet | Show no text snippet or video preview. |
indexifembedded | Allow indexing of the content when it is embedded in another page through an iframe or similar, even though the page carries noindex. It only takes effect together with noindex. |
max-snippet: [number] | Limit the text snippet to that many characters. 0 means none, -1 means no limit. |
max-image-preview: [setting] | none, standard or large. |
max-video-preview: [number] | Limit video previews to that many seconds. 0 allows a static image only, -1 is unlimited. |
notranslate | Do not offer translation of the page in results. |
noimageindex | Do not index images on the page. |
unavailable_after: [date] | Stop showing the page after the date. RFC 822, RFC 850 and ISO 8601 formats are accepted. |
Other search engines document their own sets. Check each vendor before relying on a directive outside this list.
When to use it
- Files with no HTML head. PDFs, Word and Excel downloads, images, and video. A meta tag cannot exist in them.
- Staging, preview and internal hosts. Send
noindex, nofollowfor every response from the host. It is quicker to apply than editing templates. - Pagination, filter and search-result URLs that you want crawled for links but not indexed:
noindex, follow. - Expiring content.
unavailable_after: 31 Dec 2027 23:59:59 GMTfor a promotion page.
If the content is confidential, use authentication. noindex is a request to crawlers, not access control.
Configuration
nginx
# Every PDF and Word document on the site
location ~* \.(pdf|docx?|xlsx?)$ {
add_header X-Robots-Tag "noindex, nofollow" always;
try_files $uri =404;
}
# Whole staging host
server {
server_name staging.example.com;
add_header X-Robots-Tag "noindex, nofollow" always;
}
Two nginx behaviours cause most “it works on one URL and not another” bugs. Without always, add_header only applies to 200, 201, 204, 206 and the 3xx codes, so a 404 or 500 page goes out without it. And add_header inherits from the enclosing level only if the current level has no add_header of its own: put a Cache-Control add_header in a location and the server-level X-Robots-Tag disappears for that location. nginx 1.29.3 added add_header_inherit merge to change that. On older versions, repeat the header in each block.
Apache
<FilesMatch "\.(pdf|docx?|xlsx?)$">
Header set X-Robots-Tag "noindex, nofollow"
</FilesMatch>
This is Google’s own example pattern for PDFs and requires mod_headers. Use Header always set if error responses need it too.
Cloudflare
For a site proxied through Cloudflare, add a Response Header Transform Rule (Rules, Transform Rules, Modify Response Header): set a static header named X-Robots-Tag with value noindex, nofollow, matched by a custom filter expression such as:
ends_with(http.request.uri.path, ".pdf") or http.host eq "staging.example.com"
For Cloudflare Pages static assets, use a _headers file in the build output. Cloudflare documents X-Robots-Tag: noindex for preview deployments as a use case:
https://:project.pages.dev/*
X-Robots-Tag: noindex
_headers rules do not apply to responses generated by Pages Functions. Set the header in the Function code for those.
Express
app.use('/downloads', (req, res, next) => {
res.set('X-Robots-Tag', 'noindex, nofollow')
next()
})
The robots.txt trap
robots.txt controls crawling. X-Robots-Tag is read from the response of a crawl. If a URL is disallowed, the crawler never makes the request, so it never sees the header. Google’s documentation says it plainly: if the page is blocked by robots.txt the crawler never sees the noindex rule and the page can still appear in search results. Google also does not support a noindex line inside robots.txt.
So the sequence for removing something is:
- Remove the Disallow rule for that path.
- Serve
X-Robots-Tag: noindexon the URL. - Request a recrawl in Search Console’s URL Inspection tool, or wait.
- Only after it has dropped out, if you want to stop crawling, add the Disallow back.
Verify
curl -sI https://example.com/files/report.pdf | grep -i x-robots-tag
Run it against the canonical, final URL. Then run it against a 404 path on the same host to see whether the header survives error responses. Search Console’s URL Inspection shows whether Google saw noindex on its last fetch.
Related
- Content-Type, the header that tells crawlers a URL is a PDF or an image in the first place
- Cache-Control: crawlers recrawl on their own schedule, but a long
max-ageat your CDN can keep serving an old header
Frequently asked questions
What is the X-Robots-Tag header?
It is a response header that gives crawlers the same indexing instructions as the robots meta tag, such as noindex or nofollow. Because it travels in the HTTP response, it works for files that have no HTML head to put a meta tag in: PDFs, images, video, and API responses.
Why is my X-Robots-Tag noindex not working?
The usual cause is robots.txt. If the URL is disallowed, Googlebot never fetches it and never sees the header, so the page can stay in the index as a URL-only result. Remove the Disallow, let it be crawled, and wait for a recrawl. Other causes are an add_header in a nested nginx block that silently dropped the inherited header, a CDN that strips or overwrites it, or a redirect that sends the crawler to a different URL.
Should I use X-Robots-Tag or robots.txt?
They answer different questions. robots.txt controls crawling, and Google does not support noindex inside it. X-Robots-Tag controls indexing and serving, but only for URLs the crawler is allowed to fetch. To remove something from search, allow the crawl and send noindex. To save crawl budget on a URL you do not care about, use Disallow.
Is X-Robots-Tag case sensitive?
No. Google documents that the header name, the user agent name, and the directive values are not case sensitive, so NoIndex and noindex are equivalent.
What happens if the robots meta tag and X-Robots-Tag disagree?
When robots rules conflict, Google applies the more restrictive one. A page with an X-Robots-Tag of noindex and a meta tag saying index is not indexed.
Does X-Robots-Tag stop people from accessing the file?
No. It only asks well-behaved crawlers not to list the URL in search. The file remains publicly fetchable by anyone who has the link. Use authentication for anything confidential.
Sources
- Google Search Central: Robots meta tag, data-nosnippet, and X-Robots-Tag specificationsdevelopers.google.com
- Google Search Central: Block Search indexing with noindexdevelopers.google.com
- RFC 9309: Robots Exclusion Protocolrfc-editor.org
- nginx: ngx_http_headers_module add_headernginx.org
- Cloudflare Pages: Headersdevelopers.cloudflare.com
Related
Content-Type Header: Values, Examples, charset
Content-Type tells the receiver what the body is: application/json, text/html; charset=utf-8, multipart/form-data. Common values, examples and 415 fixes.
Link Header
Learn how the Link header provides resource hints and enables preloading of CSS, fonts, and scripts to improve page load performance and user experience.
Cache-Control Header: Directives, Examples and CDN Behavior
Cache-Control directives explained: max-age, no-cache vs no-store, s-maxage, stale-while-revalidate, immutable, with nginx, Cloudflare and Next.js examples.
Accept-Ranges Header
Learn how the Accept-Ranges header tells clients whether your server supports partial content requests (byte ranges) for efficient downloads and streaming.