#robots.txt
robots.txt is a plain text file at the root of a website that tells crawlers which paths they may request. Articles tagged here explain its syntax, how rules are matched and which directives search engines actually support. They show common mistakes, such as blocking CSS, JavaScript or whole sections by accident. Each guide explains the difference between blocking crawling and preventing indexing. You will also find safe templates for WordPress, shops and staging sites. A careful robots.txt protects crawl time without hiding pages that should rank.
Oct 1, 2026 · 8 min readPay-Per-Crawl and AI Licensing: Options for Site Owners
Beyond allow or block: pay-per-crawl, licensing terms and usage preferences for AI crawlers. What exists, how mature it is and what site owners can…
Oct 1, 2026 · 8 min readCommon Crawl and CCBot: What It Means for Your Website
What Common Crawl is, what its CCBot crawler collects, how the data is used for AI training and research, and how to decide whether…
Sep 27, 2026 · 8 min readShould You Let AI Train on Your Content? A Decision Guide
A practical guide to deciding whether AI companies may train on your website content: benefits, risks, what blocking can and cannot do, and how…
Sep 24, 2026 · 7 min readnoai and noimageai Meta Tags: Do They Actually Work?
What the noai and noimageai meta tags are, which systems respect them, why they are not a standard, and what actually controls AI use…
Sep 13, 2026 · 8 min readhreflang Conflicts with noindex, robots.txt and Redirects
hreflang only works between pages that can be crawled and indexed. How noindex, robots.txt, redirects and errors break clusters, and how to resolve each.
Sep 6, 2026 · 7 min readStaging Site Indexed by Google? How to Fix and Prevent It
What to do when a staging or development site shows up in Google, how to remove it safely, and how to stop staging settings…
Aug 24, 2026 · 7 min readIs Your CDN or Firewall Blocking AI Crawlers? How to Check
Robots.txt may allow AI crawlers while your CDN, firewall or security plugin blocks them. How to find hidden blocks and fix them without inviting…
Aug 13, 2026 · 7 min readGoogle-Extended Explained: What Blocking It Does and Doesn’t
Google-Extended is a robots.txt token, not a crawler. Learn what it controls in Gemini, why it does not affect Search or AI Overviews, and…
Aug 6, 2026 · 7 min readNoindex vs Disallow: How to Keep Pages Out of Search
Noindex and robots.txt Disallow solve different problems. Learn which removes pages from search, which saves crawl time, and why combining them backfires.
Aug 4, 2026 · 8 min readHow to Allow or Block AI Crawlers in robots.txt
A practical guide to AI crawler rules in robots.txt: which bots train models, which power AI search, and copy-ready patterns to allow or block…
Aug 1, 2026 · 8 min readRobots.txt for SEO: What to Block and What to Leave Open
A practical robots.txt guide: how rules are matched, what to block, what never to block, and how to test the file before one line…