Robots.txt in 2026: Advanced Rules, Wildcards and AI Crawler Control
Robots.txt is a short file with outsized consequences. Here is how rule precedence, wildcards and AI crawler tokens really work, and how to control crawlers without hurting your visibility.
A single misplaced line in robots.txt can remove an entire section of your site from search, or quietly keep AI search engines from ever seeing your best content. In 2026 the file does two jobs: it manages how Googlebot spends its crawl budget, and it sets the rules for a growing list of AI crawlers. This guide covers the syntax and logic advanced SEOs need to get both right.
What Robots.txt Does and Does Not Do
Robots.txt controls crawling, not indexing. A URL blocked in robots.txt can still appear in search results if other pages link to it, usually without a description, because Google was not allowed to read the page. Google also stopped supporting the noindex directive inside robots.txt in 2019. To keep a page out of search, leave it crawlable and use a noindex meta tag or an X-Robots-Tag header. If you block the page, Google never sees the tag.
The file also only works on crawlers that choose to obey it. Reputable bots do. Scrapers and impersonators do not, so use rate limits or a firewall for those.
Syntax Advanced Users Need to Know
- Groups: Each group starts with one or more
User-agentlines followed by rules. A crawler follows only the most specific group that matches its name. AGooglebotgroup does not inherit rules from theUser-agent: *group, so repeat any shared rules. - Wildcards:
*matches any sequence of characters and$marks the end of a URL.Disallow: /*?sessionid=blocks any URL containing that parameter, andDisallow: /*.pdf$blocks URLs ending in .pdf. - Precedence: When rules conflict, Google applies the most specific rule, meaning the one with the longest matching path. If an Allow and a Disallow match equally, the less restrictive rule wins.
- Sitemap: The
Sitemapline takes a full URL and works independently of user-agent groups. Our guide to XML sitemap architecture in 2026 covers how to structure what you point it to. - Unsupported rules: Google ignores
Crawl-delay, although some other search engines honour it. - Limits: Google reads only the first 500 KiB of a robots.txt file, and paths are case-sensitive. Each subdomain needs its own file.
This example shows precedence in action:
User-agent: *
Disallow: /cart/
Disallow: /search/
Disallow: /*?sessionid=
Allow: /search/help/
Sitemap: https://example.com/sitemap_index.xml
Here /search/help/ stays crawlable because its Allow rule is longer, and therefore more specific, than the /search/ Disallow.
Using Robots.txt to Protect Crawl Budget
Crawl budget matters mainly on large sites, or on sites that generate huge numbers of URLs from filters, sessions and internal search. Block low-value URL patterns such as faceted filter combinations, internal search results, cart and checkout paths and endless calendar pages. Do not block the CSS and JavaScript files Google needs to render your pages. Do not use robots.txt to hide duplicate content either, because Google cannot see a canonical tag on a page it cannot crawl. Use canonical tags or noindex on crawlable pages instead.
Cross-check your rules against your sitemap with the XML Sitemap Generator and Validator. A URL that appears in your sitemap but is blocked in robots.txt sends Google conflicting signals.
Managing AI Crawlers in 2026
AI companies now run separate crawlers for different jobs, and blocking the wrong one has real consequences for your visibility. They fall into three groups: training, search indexing and user-triggered fetching.
| Operator | Training | Search indexing | User-triggered |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | None | PerplexityBot | Perplexity-User |
Google-Extended (control token) | Googlebot | No separate agent | |
| Apple | Applebot-Extended | Applebot | None |
Two details trip people up. First, blocking GPTBot does not stop OpenAI's search or user-triggered agents, because they have their own tokens. Second, Google-Extended is a control token for Gemini training and grounding, not a separate crawler, and it does not affect Google Search or AI Overviews. Those depend on Googlebot, so blocking Googlebot removes you from AI features while blocking Google-Extended does not.
A common policy for sites that want AI visibility is to allow the search and citation crawlers, decide separately about training crawlers, and treat user-triggered fetchers as ordinary visitors:
# Allow AI search and citation crawlers
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
# Opt out of model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
User-triggered agents such as ChatGPT-User and Perplexity-User may not follow robots.txt at all, because a person asked for the page. If you need to block them, do it at your CDN or firewall. Vendors also add and rename crawlers, so check each company's documentation before you copy any list.
If your goal is to be cited in AI answers, blocking the search crawlers works against you. Read our guide to what GEO is and how to optimise for it, then use the AI Overview Tracker to see whether your content is showing up.
Robots.txt, Meta Robots and X-Robots-Tag
Use each control for the job it was built for. Robots.txt manages crawling across paths and URL patterns. A meta robots tag controls indexing and snippet behaviour for an individual HTML page. The X-Robots-Tag HTTP header does the same for non-HTML files such as PDFs and images, and can be applied at server level to whole directories. A page that must stay out of search needs a noindex signal that Google can crawl, while a section you simply do not want crawled needs a robots.txt rule. Mixing the two on the same URL is how pages end up stuck in the index.
Testing and Monitoring
Test every change before it ships. The robots.txt report in Search Console shows how Google fetched your file, and the URL Inspection tool shows whether a specific URL is blocked. Watch your server logs for the user-agents you have allowed or blocked, since some scrapers impersonate real crawlers. Keep the file in version control so you can roll back quickly, and re-check it after every migration, because a staging Disallow: / that reaches production is one of the most expensive SEO mistakes there is.
Common Robots.txt Mistakes
- Blocking CSS or JavaScript files that Google needs to render the page.
- Using robots.txt to deindex a page, which stops Google seeing the noindex tag.
- Leaving a staging
Disallow: /in place after launch. - Copying a 2023 AI bot list that misses today's search and user-triggered crawlers.
- Forgetting that each subdomain needs its own robots.txt file.
Where to Go Next
Robots.txt works best alongside a clean sitemap and a site you monitor closely. Read our guide to XML sitemap architecture next, and keep our Google update traffic drop audit handy for the day a crawl or ranking problem appears.
Frequently Asked Questions
Check your sitemap and crawl signals with SEOrobin's free XML Sitemap Generator and Validator.
Open the XML Sitemap Tool