Definition
robots.txt — robots.txt is a text file at the root of a site that tells crawlers which sections they may or may not crawl.
Quick answer
robots.txt controls crawling, not indexing. To keep a page out of the index use noindex — and the page must remain crawlable.
How it works
The file lives at /robots.txt and applies only to its own host and protocol. Rules are grouped by User-agent; when several rules match, the longest (most specific) wins, and Allow wins ties. The behaviour is standardised in RFC 9309.
User-agent: *
Disallow: /admin/
Disallow: /*?sort=
Sitemap: https://garbuz.online/sitemap.xmlCommon mistakes
- Disallow: / left in place after launching from staging;
- blocking CSS/JS needed for rendering;
- blocking a page that carries noindex (the crawler can no longer see the directive);
- no Sitemap directive.
How to check
Use the robots.txt report in Search Console and the robots.txt tester.
Sources
- RFC 9309: Robots Exclusion Protocol — IETF
- Introduction to robots.txt — Google Search Central