"Blocked by robots.txt": what it means and how to fix it
Your robots.txt file stops Google from crawling the page.
What it means
A Disallow rule in your robots.txt file matches this URL, so Google does not crawl it. Google can still index a blocked URL if other pages link to it, but without its content.
Common causes
- A broad
Disallowrule matches more URLs than intended. - A staging
robots.txtstayed in production.
How to fix it
- If you want the page on Google, change the
robots.txtrule so that it does not match the URL. - If the page should stay out of Google, remove it from the sitemap.
When you can ignore it
The URL is an admin, search, or filter URL that you blocked on purpose. Remove it from the sitemap.
How to find these pages
In Search Console, open Indexing > Pages and select the "Blocked by robots.txt" row. Search Console shows up to 1,000 example URLs. CrawlCoach checks up to 10,000 URLs from your sitemap for each site and groups them by cause, so you fix the most valuable pages first.