Newest first · RSS

Field notes

Rule changes at AI engines and platforms, experiments measured on live sites, and release notes.

Serving markdown to AI crawlers: 3 of 6 docs sites already doCrawling & Indexing

Serving markdown to AI crawlers: 3 of 6 docs sites already do

Anthropic, Next.js and Cloudflare return text/markdown when you append .md to a docs URL. One site returns HTML from its .md path while reporting 200 — check the content type, not the status code.

Aug 15, 202629211533 min read
Cloudflare AI Crawl Control: per-crawler visibility you may already haveCrawling & Indexing

Cloudflare AI Crawl Control: per-crawler visibility you may already have

Cloudflare AI Crawl Control shows which AI services fetched your pages and lets you allow or block each one. Available on all plans, zero configuration — and most site owners have never opened it.

Aug 15, 20261386913 min read
Naming a crawler usually means blocking it: 82 of 99 groupsCrawling & Indexing

Naming a crawler usually means blocking it: 82 of 99 groups

Of 99 named AI crawler groups across 14 sites, 82 were a full-site Disallow. But three sites named crawlers and blocked none — counting names conflates two opposite decisions.

Aug 15, 202627931893 min read
Content signals in robots.txt: 7 of 34 sites have themCrawling & Indexing

Content signals in robots.txt: 7 of 34 sites have them

Content signals in robots.txt state what may be done with your content after it is fetched. Cloudflare launched them in September 2025 and applied them to 3.8 million domains. We found them on 7 of 34 sites.

Aug 15, 202615151143 min read
Most homepages carry no structured data: 9 of 14Crawling & Indexing

Most homepages carry no structured data: 9 of 14

Of 14 well-known homepages fetched as an AI crawler, 9 delivered zero JSON-LD blocks. The bar is lower than the volume of schema advice suggests.

Aug 15, 202629741523 min read
One company, several crawlers: blocking the wrong one costs youCrawling & Indexing

One company, several crawlers: blocking the wrong one costs you

Every major AI company now runs several crawlers with different jobs. Blocking GPTBot does not remove you from ChatGPT — that is OAI-SearchBot. Here is the map from each operator's own docs.

Aug 15, 202623431683 min read
Your docs are emptier than your homepage: 7 of 10 sitesCrawling & Indexing

Your docs are emptier than your homepage: 7 of 10 sites

On 7 of 10 well-known sites, the documentation page delivered less readable text to an AI crawler than the marketing homepage. On linear.app the docs returned an eighth as much.

Aug 15, 20261253694 min read
There is no noai directive — nosnippet is the real AI leverCrawling & Indexing

There is no noai directive — nosnippet is the real AI lever

There is no noai directive in Google's documentation. The tag that controls what AI Overviews and AI Mode may quote is max-snippet, and Google names those surfaces explicitly.

Aug 15, 20261126574 min read
Schema markup for AI search: which types are still liveCrawling & Indexing

Schema markup for AI search: which types are still live

Schema markup for AI search is worth adding, but two of the most-recommended types no longer produce rich results. Here is what Google still supports, four blocks to paste, and the policy line that gets sites penalised.

Aug 15, 202624851656 min read
Who blocks AI crawlers and who publishes llms.txt: 30 sites, two opposite strategiesCrawling & Indexing

Who blocks AI crawlers and who publishes llms.txt: 30 sites, two opposite strategies

We fetched robots.txt and llms.txt from 30 well-known sites. News publishers name AI crawlers by the dozen and publish no llms.txt; developer-tool companies do the exact opposite.

Aug 15, 202629832374 min read
llms.txt: what it is, a template, and what it cannot doCrawling & Indexing

llms.txt: what it is, a template, and what it cannot do

An llms.txt lists your most useful pages so an agent can find them without crawling everything. Here is the format, a template, who actually publishes one, and the five things it does not do.

Aug 15, 202615131197 min read
What AI crawlers get from JavaScript sites: 342 to 25,082 charactersCrawling & Indexing

What AI crawlers get from JavaScript sites: 342 to 25,082 characters

What AI crawlers get from JavaScript sites varies 73-fold. We fetched 15 well-known homepages as an AI crawler: 14 returned 200, and the readable text ranged from 342 to 25,082 characters.

Aug 15, 20261348744 min read
1 / 2
Go to