What is robots.txt? It is one of the smallest files on your entire website, often just a few lines of plain text, yet it quietly controls one of the most important relationships your site has: the one with search engine crawlers. Get it wrong, and you can accidentally block Google from your entire site overnight. Get it right, and you steer crawlers toward the pages that matter most.
This guide covers exactly what a robots.txt file is, how it works, the syntax you need to know, real examples you can copy, and how to audit yours for mistakes, including how to handle the newer wave of AI crawlers in 2026.
Robots.txt is a plain-text file placed at the root of a website (yourdomain.com/robots.txt) that tells search engine and AI crawlers which parts of the site they are allowed or not allowed to crawl. It manages crawler traffic and crawl budget, not indexing: a page blocked in robots.txt can still appear in Google’s index if other pages link to it. To stop a page from appearing in search results entirely, use a noindex meta tag instead.
⚡ Key Takeaways
- Robots.txt controls crawling, not indexing. Blocking a URL in robots.txt does not guarantee it stays out of Google’s search results.
- The file must live at the exact root of your domain (yourdomain.com/robots.txt) to be recognized by crawlers.
- Every website benefits from having a robots.txt file, even a simple one, to prevent wasted crawl activity and give crawlers clear direction.
- Never block CSS or JavaScript files; Google needs them to render and understand your pages properly.
- Robots.txt is publicly viewable by anyone, so it should never be used to hide sensitive or private URLs.
- In 2026, robots.txt is also the primary tool for controlling access by AI crawlers such as GPTBot, ClaudeBot, and PerplexityBot.
- Google Search Console’s robots.txt report is the fastest way to confirm your file is valid and doing what you intend.
What Is a Robots.txt File?
A robots.txt file is a plain-text file that lives in the root directory of a website and follows the Robots Exclusion Protocol, a long-standing, informal standard that well-behaved crawlers agree to respect. It is the first thing most search engine bots check before crawling anything else on your domain.
Think of it as a set of instructions posted at the entrance of your website. It does not lock any doors. It simply tells visiting crawlers which rooms they are welcome to enter and which ones you would rather they skip. Reputable crawlers, like Googlebot and Bingbot, read and respect these instructions. Malicious bots and scrapers are free to ignore them entirely, which is why robots.txt should never be treated as a security measure.
💡 Quick definition: Robots.txt is a permissions file for crawlers, not a lock. It is a request, not an enforcement mechanism.
What Does Robots.txt Actually Do?
This is the single most misunderstood part of robots.txt, so it is worth being precise. Robots.txt controls crawling, the process of a bot visiting and reading a page. It does not directly control indexing, the process of that page being stored and shown in search results.
If a page is disallowed in robots.txt but other websites or internal pages link to it, Google can still index the URL, sometimes showing it in search results with a description like “No information is available for this page.” Google simply cannot read the page’s content because it never crawled it, but the URL itself is still eligible to appear.
- Robots.txt controls: whether a crawler is allowed to request and read a URL at all.
- Robots.txt does NOT control: whether that URL can appear in search results, how it ranks, or whether it gets indexed.
If your goal is to stop a page from ever appearing in search results, the correct tool is a noindex meta tag or X-Robots-Tag HTTP header, covered in detail below, not robots.txt.
Do You Even Need a Robots.txt File?
Yes, virtually every website benefits from having one, even a very small site. Without a robots.txt file, crawlers will simply crawl everything they can find, which is not necessarily a problem for a five-page brochure site but becomes wasteful and messy as a site grows.
| Reason to Have One | Why It Matters |
|---|---|
| Prevents crawl waste | Keeps bots from spending time on admin pages, internal search results, and duplicate parameter URLs |
| Protects server resources | Reduces unnecessary bot traffic on large or resource-constrained sites |
| Points to your sitemap | The Sitemap directive helps crawlers discover your XML sitemap immediately |
| Gives you crawler-level control | Lets you allow helpful bots while blocking unwanted scrapers or AI crawlers |
| Prevents accidental full blocks | An intentional, reviewed file is safer than no file at all combined with a misconfigured CMS default |
The only sites that can reasonably skip it are extremely small, static sites with nothing to hide from crawlers and no sitemap to reference, and even then, having a minimal file that simply allows everything is considered best practice.
How Robots.txt Works: Syntax and Directives
A robots.txt file is made up of one or more groups of rules. Each group starts with a User-agent line specifying which crawler the rules apply to, followed by Allow and Disallow directives.
| Directive | Purpose | Example |
|---|---|---|
User-agent | Specifies which crawler the following rules apply to | User-agent: Googlebot |
Disallow | Tells the specified crawler not to access a path | Disallow: /wp-admin/ |
Allow | Overrides a Disallow rule for a specific subpath | Allow: /wp-admin/admin-ajax.php |
Sitemap | Points crawlers to your XML sitemap location | Sitemap: https://example.com/sitemap.xml |
Crawl-delay | Requests a delay between crawl requests (ignored by Googlebot) | Crawl-delay: 10 |
Two wildcard characters give you finer control:
*matches any sequence of characters, useful for pattern-blocking entire types of URLs, such asDisallow: /*?sort=.$anchors a rule to the end of a URL, useful for blocking a specific file type, such asDisallow: /*.pdf$.
Rules are also case-sensitive, and the most specific matching rule wins when an Allow and Disallow conflict for the same crawler.
Robots.txt vs. Noindex vs. X-Robots-Tag
These three tools are frequently confused, but they solve different problems and are sometimes used together incorrectly, which cancels out the intended effect entirely.
| Tool | Where It Lives | What It Controls | Best Used For |
|---|---|---|---|
| Robots.txt | Root file (/robots.txt) | Whether a crawler may request the page at all | Managing crawl budget, blocking entire sections or bots |
| Meta robots (noindex) | Inside the page’s <head> | Whether the page is allowed to appear in search results | Removing individual pages from search results while still allowing them to be crawled |
| X-Robots-Tag | HTTP response header | Same as noindex, but works for non-HTML files (PDFs, images) | Keeping PDFs, images, or other files out of search results |
⚠️ Common trap: If a page is both disallowed in robots.txt AND has a noindex tag, Google can never crawl the page to see the noindex tag in the first place, so the page can still be indexed from external links. To fully remove a page, allow crawling and use noindex; do not disallow it.
Where to Find (or Put) Your Robots.txt File
Every robots.txt file must sit at the exact root of the domain, never in a subfolder, for crawlers to find it. You can check any website’s file instantly by adding /robots.txt to its homepage URL, for example https://www.1solutions.biz/robots.txt.
- Correct:
https://example.com/robots.txt - Incorrect:
https://example.com/blog/robots.txt - Each subdomain needs its own separate robots.txt file, since crawlers treat subdomains as distinct hosts.
On WordPress, if no physical file exists, WordPress generates a virtual one automatically. Most SEO plugins, including Yoast SEO and Rank Math, let you edit its contents directly from the plugin settings without needing FTP access. Reviewing it belongs on any recurring maintenance routine; see our WordPress maintenance checklist for the other technical items worth checking alongside it.
How to Create a Robots.txt File
Step 1: Decide What You Actually Need to Block
List the sections of your site that provide no value to crawlers: admin areas, internal search result pages, staging folders, cart and checkout steps, and duplicate parameter-driven URLs.
Step 2: Write the File in Plain Text
Create a file named exactly robots.txt using any plain-text editor. Never use a word processor, which can insert hidden formatting characters that break the file.
Step 3: Add Your Sitemap Reference
Always include a Sitemap line pointing to your XML sitemap. This is one of the easiest, highest-value lines in the entire file.
Step 4: Upload It to Your Root Directory
Upload via FTP to your site’s root, or, on WordPress, edit it directly through Yoast SEO (Tools → File Editor) or Rank Math (General Settings → Edit robots.txt) without touching server files at all.
Step 5: Test Before You Trust It
Never assume a hand-written file works as intended. Test it using the methods in the auditing section below before considering the job done.
Robots.txt Best Practices
- Never block CSS or JavaScript files. Google renders pages like a browser does, and blocking these resources can cause Google to misjudge your page’s layout, mobile-friendliness, and content quality.
- Never use it to hide sensitive information. Robots.txt files are fully public. Listing a private admin path in it can actually advertise that path’s existence to anyone who checks.
- Keep it as simple as possible. Overly complex rule sets are where mistakes hide; every line should have a clear, intentional reason for existing.
- One rule per line. Combining multiple paths on a single Disallow line is not supported by the standard and will not work as expected.
- Always include your sitemap. It costs one line and gives crawlers a direct path to your most important URLs.
- Re-check it after every major site migration or CMS change. Migrations are the single most common cause of accidental full-site blocks; see our guide on how to redesign a website without losing SEO for the full pre-launch checklist this fits into.
🚨 The most expensive mistake in SEO: A single stray line, Disallow: /, blocks your entire website from every crawler it applies to. This is one of the most common causes of a sudden, site-wide traffic collapse after a redesign or migration.
Blocking AI Crawlers (2026 Guide)
Robots.txt has taken on a new role since the rise of AI-powered search and large language models. A growing list of AI companies now operate crawlers that scan the web to train models or generate real-time answers, and most respect robots.txt directives the same way traditional search engines do.
| Crawler | Operated By | Purpose |
|---|---|---|
| GPTBot | OpenAI | Model training |
| ChatGPT-User | OpenAI | Real-time browsing during a ChatGPT session |
| ClaudeBot | Anthropic | Model training |
| Claude-User | Anthropic | Real-time browsing during a Claude session |
| PerplexityBot | Perplexity AI | Real-time answer generation |
| CCBot | Common Crawl | Open dataset used by many AI labs for training |
| Google-Extended | Controls use of your content for Gemini and AI features, separate from normal Googlebot search indexing |
Whether to block these is a business decision, not a technical one. Blocking training crawlers (GPTBot, ClaudeBot, CCBot) can reduce your content’s chance of being cited in AI answers, but keeps your content from being used to train models without direct benefit to you. Blocking real-time browsing crawlers (ChatGPT-User, Claude-User, PerplexityBot) can make your pages invisible to users asking AI assistants questions your content answers, cutting off a growing discovery channel. This decision is really an extension of your broader answer engine optimization strategy, and if Perplexity specifically matters to your traffic mix, our Perplexity SEO guide covers how PerplexityBot fits into that decision in more depth.
# Example: allow AI answer engines, block AI training-only crawlers
User-agent: GPTBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-User
Allow: /
For most content-driven businesses looking to grow visibility in AI Overviews and AI search assistants, keeping these crawlers allowed is the more common approach in 2026. Blocking them is more defensible for proprietary data, paywalled content, or publishers whose business model depends on direct traffic rather than citations.
Example Robots.txt Files
Allow Everything (Minimal Site)
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Standard WordPress Site
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-includes/
Disallow: /?s=
Disallow: /search/
Sitemap: https://example.com/sitemap_index.xml
Ecommerce Site with Filter Parameters
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Disallow: /*?sort=
Disallow: /*?filter=
Disallow: /*?color=
Sitemap: https://example.com/sitemap.xml
Blocking Everything (Staging Sites Only)
User-agent: *
Disallow: /
💡 Pro Tip: This last example is the exact rule that causes disasters when a staging site’s robots.txt accidentally gets pushed to production during a launch. Always double-check this file immediately after any site migration.
How to Test and Audit Your Robots.txt File
- Google Search Console: Use the URL Inspection tool on individual pages to see whether Google considers them blocked by robots.txt, and check the Settings > Crawling report for how often Google fetches your file.
- Manual check: Visit yourdomain.com/robots.txt directly in a browser to confirm the file loads, returns a 200 status, and contains what you expect.
- Validate syntax: Check for typos in directive names, missing colons, and paths that do not start with a forward slash, all common hand-editing errors.
- Check for orphaned rules: Remove Disallow rules for pages, folders, or plugins that no longer exist on your site.
- Confirm your sitemap line resolves: A broken or outdated Sitemap URL in robots.txt wastes an easy discovery opportunity for crawlers.
Robots.txt and Crawl Budget
Robots.txt is one of the primary levers for managing crawl budget, the finite number of URLs Google will crawl on your site within a given period. By disallowing low-value, duplicate, or parameter-heavy URLs, you free up crawl budget for the pages that actually drive rankings and revenue. Large or frequently updated sites see the biggest impact here; if you manage an ecommerce store or a content-heavy publication, read our full guide on crawl budget optimization for the complete step-by-step process.
Robots.txt Checklist
- ☐ File is located at the exact root (yourdomain.com/robots.txt)
- ☐ Returns a 200 status code, not a 404 or redirect
- ☐ No accidental
Disallow: /for User-agent: * - ☐ CSS and JavaScript files are not blocked
- ☐ Sitemap directive is present and points to a working sitemap
- ☐ No sensitive or private paths are listed
- ☐ AI crawler rules reflect an intentional decision, not a default
- ☐ File has been checked in Google Search Console after the last update
Frequently Asked Questions
Is robots.txt required for every website?
It is not strictly required, but it is considered best practice for every website regardless of size. Without one, crawlers will attempt to crawl everything they can find, which can waste crawl budget and, in the worst case, expose URLs you never intended to be crawled at all.
Does robots.txt guarantee a page stays out of Google?
No. Robots.txt only prevents crawling. A disallowed URL can still be indexed and shown in search results if other pages link to it, just without a description, since Google never crawled the content. To guarantee exclusion from search results, use a noindex tag instead.
Can robots.txt hurt my SEO?
Yes, significantly, if misconfigured. The most damaging mistake is an accidental Disallow: /, which blocks an entire site from being crawled and can cause a rapid, dramatic drop in organic traffic. Blocking CSS or JavaScript files can also hurt rankings by preventing Google from properly rendering and evaluating your pages. If you are troubleshooting a sudden traffic drop and are not sure robots.txt is the cause, our guide on why your website ranking dropped walks through the full list of likely culprits.
How often should I update my robots.txt file?
There is no fixed schedule. Update it whenever your site structure changes meaningfully, such as adding new sections, changing platforms, or launching a redesign, and always re-check it immediately after any migration or major deployment.
Do all crawlers respect robots.txt?
No. Reputable crawlers from Google, Bing, and most major AI companies respect it voluntarily. Malicious scrapers, spam bots, and some smaller or less scrupulous crawlers can and do ignore it entirely, which is why robots.txt should never be relied on to protect sensitive data.
Can I have different robots.txt rules for different bots?
Yes. Each User-agent block applies only to the crawler named in it. This lets you, for example, allow Googlebot full access while blocking a specific AI training crawler, simply by writing separate groups of rules for each.
Getting Professional Help with Technical SEO
A misconfigured robots.txt file is a small technical detail with an outsized ability to damage visibility overnight, and auditing one properly means understanding how it interacts with your sitemap, your crawl budget, and your broader indexation strategy.
At 1Solutions, our professional SEO services include full technical audits covering robots.txt configuration, crawl budget analysis, sitemap health, and indexation troubleshooting. With over 15 years of experience, we have caught and fixed robots.txt errors before they turned into full-blown traffic incidents, and built AI-crawler strategies for clients navigating the shift toward answer engines.
If you would rather this kind of file get checked on a schedule instead of after something breaks, our website maintenance services fold robots.txt and sitemap reviews into ongoing site upkeep, alongside performance, security, and update monitoring.
Conclusion
What is robots.txt? It is a small, plain-text file with outsized responsibility: directing how search engines and AI crawlers experience your website. Used correctly, it protects crawl budget, keeps low-value pages out of the crawl queue, and gives you deliberate control over the newest generation of AI crawlers. Used carelessly, it can take your entire site out of Google overnight. Treat it as a file worth reviewing on a schedule, not a set-and-forget setting from your CMS defaults.




