← back to the board

What AI crawlers actually fetch

Plenty of people will tell you which crawlers to allow. Almost nobody publishes what the crawlers then did, because almost nobody keeps the log. This is one small site's server data, published every day: no sampling, no estimates, every request that identified itself as a known crawler.

Every page here is served at two addresses — the HTML one, and a .md twin at the same path. The share of crawler requests that take the markdown is the number worth watching, because it is a measurement of whether offering a plain-text version is worth the trouble, and it is not a number anyone can look up.

No day has been reported yet. The first report is built the morning after the first full day of logging.