Direct answer
Robots.txt controls whether a compliant crawler may request a URL. A meta robots noindex directive or X-Robots-Tag asks that a crawlable response not be indexed. Authentication controls who can access private content. Noindex is not a valid robots.txt rule, and blocking a URL can prevent a crawler from seeing its noindex directive.
Choose the control that matches the outcome
| Required outcome | Primary control | Important boundary |
|---|---|---|
| Keep private content unavailable | Authentication and authorization | Robots.txt and noindex do not secure content |
| Reduce crawling of a public URL space | Robots.txt | A blocked crawler may still know the URL and cannot read page-level noindex |
| Allow crawling but request exclusion from search results | Meta robots noindex or X-Robots-Tag | The crawler must be allowed to fetch the response and see the directive |
| Remove a page that no longer exists | 404 or 410 | Do not keep a false 200 page or redirect unrelated URLs home |
| Move a page permanently | 301 or 308 redirect | Update internal links and avoid chains |
Valid implementation examples
A robots.txt rule belongs in the host-level robots file. A noindex directive belongs in a crawlable HTML response or HTTP header. Writing noindex inside robots.txt is not a valid way to request index exclusion.
User-agent: *
Disallow: /internal-preview/<meta name="robots" content="noindex, follow">X-Robots-Tag: noindex, followCommon scenarios
- Staging environment: protect it with authentication or network access control; do not rely on robots.txt.
- Internal search results: keep navigation useful, use noindex when appropriate and prevent unbounded parameter discovery by design.
- PDF that should remain available but not indexed: send an X-Robots-Tag noindex header and allow the crawler to fetch it.
- Large faceted space that should not be crawled: first remove uncontrolled links and traps, then use robots rules only where blocking fits the verified design.
- Already indexed URL: do not block it before the crawler can observe noindex or a truthful removal status.
Deployment verification
- Request the exact URL and record status, redirect chain, response headers and final HTML.
- Test the matching robots rule for the intended crawler and host.
- Confirm a noindex directive is visible in the final crawlable response and is not contradicted by a redirect or canonical plan.
- Check sitemaps and internal links so discovery signals match the intended lifecycle.
- Monitor Search Console after recrawling; removal timing is controlled by the search engine, not by a fixed promise.
Common mistakes
- Putting noindex in robots.txt.
- Blocking a page before Google can see noindex.
- Treating robots as security.
- Assuming removal is immediate.
Limitations
- Crawler support varies.
- Index removal timing is outside the publisher's control.
- A directive does not replace truthful status and lifecycle handling.
Primary sources
- Introduction to robots.txt — Google Search Central; current 2026-08-19; 2026-08-19.
- Robots meta tag, data-nosnippet, and X-Robots-Tag specifications — Google Search Central; current 2026-08-19; 2026-08-19.
- Robots Exclusion Protocol — IETF; 2022-09; 2026-08-14.