# Known-Good > An index of websites verified to be usable by AI agents. We read every robots.txt in a > full public crawl of the web, find the sites that invite agents in, then probe them live > to check whether they built anything an agent can actually use. About half had. Known-Good grades sites against a published rubric and dates every claim. A declaration is not a capability: we test rather than trust the label. Verification of the full index is in progress; this site is v0.1. We declare `ai-input=yes` and `ai-train=no`. Agents are welcome to read this site and use it to answer someone's question. Please do not use it as training data. ## Pages - [Homepage](https://knowngood.sh/index.md): what we found, and the three things that make a site agent-ready. Markdown twin of the homepage. - [Crawler information](https://knowngood.sh/bot.md): what our KnownGood-Verifier crawler does, what it fetches, and how to block it. Markdown twin of /bot. ## Machine endpoints - [Crawler IP ranges](https://knowngood.sh/bot/ips.json): the published source addresses our verifier crawls from, so you can confirm a request is genuinely ours. - [robots.txt](https://knowngood.sh/robots.txt): our own content signals, declared honestly. ## Key findings - 83,900,812 robots.txt files read across a full public crawl of the web. - 2,171,626 of them carry a content signal — a declared stance on AI use. - 57,740 domains explicitly welcome AI agents to read them. That is roughly 1 in 1,453. - About 1 in 2 of the sites that decided for themselves have built something an agent can use. - 86% of all content-signal declarations are one CDN's default rather than a decision anyone made. ## Contact - Free 30-minute agent-readiness review: https://cal.com/knowngood.sh/30min - Crawler questions or opt-out: bot@knowngood.sh - Everything else: hello@knowngood.sh ## Not yet live Content negotiation on the `Accept` header, the search API, and an MCP server arrive with v1. We do not list endpoints that do not exist.